It's been so long since I wrote a new piece. That's due to the hectic lifestyle I lived since I got married and my wonderful daughter is born. Juggling between fulfilling my work duties, being a loving husband and a doting father, it's been a struggle. Even my personal downtime is also compromised.
But recently, with the help of AI, I'm starting to reclaim back my time. It started with Google's Gemini Chat. To be accurate, Gemini wasn't the first LLM I used, but it's the first LLM that I use extensively for productivity. ChatGPT was the first, but it was more like exploratory usage.
Initially, Gemini was merely a tool to learn something really fast. Instead of scouring the internet with search engines to look for answers, Gemini could give me the answer within 1 minute. It used to be that I had to piece the information puzzle together one by one, page by page, site by site. What took hours was done in minutes.
Eventually, I started feeding it some PHP codes. It started to give me good solutions in tiny code snippets. It was readable and understandable. After all, when the code is 10 lines, it's easy to vet through line by line within 5 minutes. But then, that's the part I have a realization. While it took me 5 minutes to fully digest the code, Gemini took less than 1 minute to generate not just the code, but also the conversation explaining the code. This was the start of me handing over my over-20 years of engineering to AI.
Next came larger codebase. I started feeding entire files to Gemini to analyse and generate code. Since Gemini limited uploading 10 files at a time, I couldn't simply drag-and-drop. Some files were way too large. I could extract code snippets manually, but it turns out that the time I take extracting code snippets takes far longer than uploading 20 files, 10 files each time.
I also happen to have a deprecated Tabnine subscription, which was simply a VSCode extension and it was just plain old chat wrapper with some Claude models. Tabnine had restrictions, so I often hit rate-limits. But still, I made use of it as much as possible by combining Gemini and Tabnine's models intelligence by facilitating both of them to argue the best way to approach problems. Then came the day I anticipated. Tabnine cancelled my deprecated subscription and refunding any remaining unused amount.
At this point, I also started a new project and, this time, Gemini built the entire app without me writing a line of code. I was still using Gemini Chat up to this point. So it was a lot of drag-and-drop and copy-pasting between Visual Studio Code and Gemini Chat. Other administrative task include creating new files and folders, and performing git actions. Even git messages were written by Gemini at this point. But not a single line of code was written by me. It worked!
It is also around at this time that OpenClaw started to gain in popularity. After reading the horrors of security vulnerabilities, I decided to approach OpenClaw with extreme caution. Till today, I never expose OpenClaw to any chat app. I've also heard about the token burn using OpenClaw, so I used my company's Gemini subscription to experiment with OpenClaw. At that time, Gemini CLI had a generous limit and it could be hooked up with OpenClaw. This is the moment I realise the capabilities of Agentic AI. This is the moment everything changed.
Agentic AI is the future!
At least for the foreseeable future. So I poured by heart and soul towards this new shiny thing: Agentic AI. I was deeply amazed by Gemini's agentic capabilities riding on OpenClaw. This time round, instead of my usual drag-and-drop to Gemini, I simply pointed OpenClaw to my codebase and described my problem. Almost instantly, my problem was solved. It even ran tests to confirm that it works.
But of course, such generosity from Google couldn't last. Google tighten the rate limits to the extent that OpenClaw was unusable. I was almost always hitting rate limits. But I wasn't about to give up so soon. I knew I must eventually pay for my usage, but I wasn't sure how expensive it's going to be. Since I'm already an AWS customer and AWS bedrock host many models, I decided to try out Nova and Claude.
Both Nova and Claude is good. It solved every problem I threw at it, though it needed some guidance with context from time to time. I guess the problems I needed solving isn't that difficult, so both models worked fine. But strangely, when the first bill came, Nova cost more than Claude despite that Nova's token costs was lower than Claude. It turns out that OpenClaw didn't use cached prompts, so every prompt was charged at the input non-cached rates.
So I switched to using Claude exclusively. I used mostly Haiku 3.5, since most of my prompts ain't some complex math. Then the 2nd bill came. I was burning ~USD100 a day. Ouch!
That's it. I had to stop using hosted models. They are fast and good, but it burned a whole in my wallet. This time round, I focused on running local models. Since I already have a RTX 4060, why not make good use of it. Turns out, easier said than done.
Models that can fit inside the limited 8GB vRAM simply ain't smart enough. It can hold a basic conversation, but that's it. Getting it to power OpenClaw is nearly impossible. Offloading some of the weights to RAM just slows the model and makes the whole system unstable. Like the whole computer gets unstable. The best I could use my GPU for is to run Ornith1.5 9B to generate git messages through Zoo Code.
So my mission is the find a way to run larger models on dedicated hardware. The largest vRAM non-Apple, consumer-grade device I could get my hands on at that time was Ryzen AI Max+ 395, costing me more than SGD4000. But hey, if I was burning USD100 a day running Claude Haiku, I would easily break even within 2 months, running to smartest open-weight model I could fit.
And there I got it. The day it arrived, I was so eager to try it. I allocated the maximum amount of RAM I could to GPU and get my hands dirty with the first LLM. Since it was my first attempt, I just wanted it working as soon as possible. Ollama was my immediate choice. It was simple and effective. Most importantly, it simply worked.
But as smooth as it was with deployment, the price of this simplicity was the prefill and inference speed. Ollama comes with a tax on every prompt. After spending this money, I wasn't about to leave any performance on the table. So I figured out how to run Llama.cpp raw. No more middleman.
Turns out to be harder, but once I get the execution commands right, I just needed a script to execute the same way every time.
These days, I'm running Qwen3.8-flash-next. It's a bit on the slow side, but it works as good as Claude Sonnet 5. I'm contemplating whether to switch to Ornith1.5 35B A3B. It is suppose to be slightly worse than Qwen3.8-flash-next, but runs at far better speeds. But for now, as long as it works and gets things right, maybe I'll stick to Qwen3.8-flash-next for the time being.
I'm also using Cursor for my development. I've got to say, the auto+composer token limit sure is generous. Of my months of subscription since May, I only ever consumed 100% of my monthly allocation once. I could use my Qwen3.8-flash-next with Cursor, but I don't see the need to. Why use my local model when I still have plenty of tokens left with Cursor?
My electricity bill has certainly gone up. I'm using at least an extra 100kWh every month, with amounts to more than SGD34 every month. But this amount is far easier to swallow than the USD100/day I was spending before. I would happily pay my utilities over AWS bedrock. The other effects of running my local model is that the GPU is heating up my room, so I am using a bit more air-conditioning than I usually would. Still, happy to pay this bill over AWS bedrock anytime.
I don't know when I can write my next post, but I guess this post is a sign that I have managed to reclaim back a little bit of my time so far. I hope I can continue to reclaim further more. I won't let AI write my posts. After all, you're here because you want to hear my voice, not some version of AI I control. But the images will be AI generated.