Posts

No more token anxiety

Image
Here's my openrouter token usage in a specific day: (It's a flex, I know) Approximately 40M ( < 5$) tokens in 4 hours. More than the weekly limit in Claude Max X20 no matter which model you use. Estimated Weekly Token Limit (Claude Max 20x - 200$ monthly subscription) Claude Opus: ~1.5 Million to 2.5 Million total tokens per week Claude Sonnet: ~3 Million to 5 Million total tokens per week Claude Haiku: ~12 Million to 20 Million+ total tokens per week For me, that's a good enough justification to move from claude to the carefree world of open router and opencode.  Using opencode + openrouter I get: - Better latency. - Extremely lower price. - No monthly subscription (pay as you go) - No token anxiety! I can burn 100M tokens in one hour if I want to. Wanna join the token maxing party, there you go: https://openrouter.ai/docs/...

How to make a side project last

Image
In the past decade or so I've started and abandoned numerous side projects so here's my biased opinion about side project maintenance - Easy deployment and minimal dependencies is the 'make or break' for side projects . The flow of most of the abandoned projects looks like this: Get excited about a shiny new idea. Build it Deploy it It runs for about a year. Oh no, a bug Getting the dev-env + deployment up and running will take an hour so I never get to actually do it! ☠️ Dead project ☠️ The ones that are still active all share one thing in common - dead simple development env + deployment mechanism. My best project is using this tech stack: Only language - golang compiled to one statically linked binary so the deployment is just scp and run strong type system - I get some degree of protection even though I don't write any tests. Dev container for the development env with a static go version. Version control - private github Server - Raspberry Pi Transport - Telegra...

Software that mutates with you (zero friction dev)

Image
  Since coding agents started there is a big shift in the way I work. Instead of thinking about the problem in terms of algorithms, I try to design an "islands" based architecture where each "island" can be filled with pure garbage but the full software stays solid by strict apis and data boundaries. A month ago our first boy was born (a bit ahead of time) and we needed to monitor his eating/pooping habits for health reasons. We looked for an off the shelf baby logging app - there were a few good options but none of them was a perfect fit. So I built one! Setup: Dev — Opencode connected to openrouter, using DeepSeek V4 Flash latest Server — Raspberry Pi Reachability — Tailscale The main "garbage islands" where: Whatsapp chat parsing Analysis page generator Measurement addition page Web server Each of these islands is pure AI slop but since the design is extremely simple, it holds pretty well. It's remarkable how fast (and cheap) software development h...

I'm using dev-containers and you should too

Image
A Dev Container is an entire, pre-configured development system encapsulated inside a Docker container. Unlike a production container that just runs compiled code, the Dev Container hosts your entire development lifecycle (coding -> testing -> debugging -> building). I used to manage a team of ~10 software developers working on multiple unrelated projects. Each project had its own tech stack and it's own undocumented build system installation procedure. Keeping up with the different tools of each project was extremely hard and I was stuck with the feeling that there must be a better way to handle this mess. Fast forward to 2022 - I joined the founding team of a small startup and it was time to put the lessons learned to the test. We wanted to create an organization where developers would be able to understand and change the entire codebase and not just their main domain. In order to do that, we needed to lower the gates between backend, frontend and any other infra in the...

Tricking our prehistoric brains to spot a bug a mile away

Image
Have you ever done something just because it was cool but it ended up being actually useful? When I started coding, The Matrix movies just came out and ever updating green text over black background was all the rage. I didn't have anything interesting to run in my terminal except vim so I ran  htop in one terminal while coding in the other.   I had a eureka moment while working on a multi-threaded application and all of a sudden the htop CPU bars were all flashing green and red. Something  felt  wrong, so I checked my recent changes and found a slow fork bomb hiding in plain sight. That's when it hit me, the same brain that was trained for millions of years to spot a snake approaching from the corner is now able to spot a bug a mile away! All I needed to do is feed those signals in the side of my screen  so my brain will do the  visual  anomaly detection in the background while I focus on the actual task. Fast forward 20 years to 2025 and I'm using ...

Can a 1.4GB LLM Be My Local Search Engine? (spoiler: yes)

Image
One of the main LLM use cases is often pitched as a search engine replacement and it got me thinking.. Can a small locally hosted LLM replace google when I'm on the go? TL;DR - Yes!  Though LLM stands for Large Language Models, some are surprisingly small. To get started, I installed Ollama , which makes tinkering with a huge number of open-source models incredibly easy. I wanted to find a model that's small enough so it can run smoothly on my laptop while being actually useful for coding. After some tinkering, I settled on qwen3:1.7b . The 1.7b version is only 1.4GB so it fits well in my laptop's RAM and it's good enough for real coding tasks! Now, I can just spin up a local LLM in my termial (!!) and ask it random stuff while coding:   For simple search tasks (and even a bit of debugging), it's as good as the full blown web based LLMs. It won't tell me today's news, but for quick code snippets, syntax and library definitions, it's perferct. (Sidenote -...

The ratchet technique - how to prevent new shit code from joining your code

Image
  There's a common problem that happens when new linters are added to an existing codebase. Let's say we have 50K rows of exeisting code and we want to add a new shiny linter that makes sure all module names are Harry Potter characters. We run the new linter for the first time and it finds 1337 linting violations . We want to prevent new code from violating these rules but we don't have the time to fix the old code  at the moment.  This is where the "Ratchet technique" comes in! We start by adding a CI merge gate workflow that counts the number of violations  (1337) and ensures the number never increases.  If someone commits new code that doesn't adhere to the new standard, the PR checks will fail. Old non-compliant code can stay bad until we have the time to fix it.  This method  halts code degredation  without any major refactoring.  Funny new behavior that was observed in the wild - developers will fix an old linter errors so they can add ...