A few months ago, I caught myself looking at a completed pull request with a lingering sense of unease.
The diff itself was clean. Twelve lines of Go in an HTTP router, a straightforward unit test, and a green badge on the CI runner. If you inspected the git commit log, the entire change looked like a ten-minute task.
What the git log failed to record was the invisible churn that produced it.
Behind those twelve lines was an autonomous agent session that ran for nearly forty minutes. It had traversed eighteen files, triggered thirty-two model turns, burned through four hundred thousand tokens, and silently stumbled down two architectural dead ends before finding the right file. I was paying the API invoice for an entire engineering afternoon, yet my tools treated the work as if it had cost nothing more than a few keystrokes.
We are living through a curious transition in software engineering. We have welcomed autonomous coding agents into our daily workflow, yet we still manage them with nineteenth-century habits and twentieth-century terminal tooling. We know the exact memory footprint of our compiled binaries, but we have zero visibility into the token burn of our daily coding sessions. We collect procedural agent skills like browser bookmarks, scattering them across dozens of vendor-specific dotfile directories until our context windows choke on alphabetical truncation.
Over the past year, I stopped treating agents as black-box appliances and started building the missing scaffolding. The goal was never to wrap them in complicated enterprise orchestration platforms. Instead, I wanted small, local-first utilities built in Go that run on my own workstation, read ordinary SQLite databases, speak standard protocols like the Model Context Protocol (MCP), and answer three basic questions: What did this session cost, what skills are my agents actually carrying, and how do we keep their work grounded in reality?
Here is the stack I run every day.
Thermal: Financial Vision from Read-Only SQLite
The first missing piece was accounting. When an agent runs in your terminal, it writes session telemetry somewhere on your local disk. OpenCode, Devin, Codex, MiMoCode, and Claude each maintain local stores. Some use JSON lines, others use SQLite, and a few store raw session state in proprietary caches.
Rather than running an intercepting proxy daemon that injects network latency or wrapping agent binaries in fragile shell shims, I built Thermal.
Thermal is an open-source terminal utility written in Go. Its operating philosophy is strictly read-only: it reaches directly into the local databases already left behind by your coding tools on disk, parses token usage and model identifiers, and calculates running streaks, daily volume, and real financial spend. Because it reads local files directly with SQLite memory-mapped I/O, generating an entire cross-tool census takes less than twenty milliseconds.
When I run thermal in my workspace today, it does not give me vague
estimates or cloud dashboard spinners. It prints an immediate, concrete
leaderboard:
THERMAL : Don't break the streak.
╭─ Token Warriors ────────────────────────────────────── tokens & spend ─╮
│ # Tool Strk Best Days Tokens Cost │
│ ───────────────────────────────────────────────────────────────────── │
│ 1. Codex 3d 7d 49d 1.0B tok ~$538.30 │
│ 2. Devin 2d 33d 45d 19.2B tok ~$20134.07 │
│ 3. OpenCode 1d 11d 30d 10.6B tok $62.91 │
│ 4. MiMoCode 1d 7d 16d 1.2B tok $56.14 │
│ 5. ZCode 1d 6d 11d 871.0M tok ~$25.01 │
│ 6. DeepSeek (DSH) 1d 1d 5d 18.0M tok ~$0.07 │
│ 7. Grok 1d 1d 1d 692.7K tok $0.44 │
│ 8. codewhale 1d 1d 3d 110.1K tok $0.02 │
│ 9. Claude 1d 1d 3d 55.5K tok ~$0.16 │
╰────────────────────────────────────────────────────────────────────────╯
╭─ Activity Hunters ────────────────────── messages & steps ─╮
│ # Tool Strk Best Days Activity │
│ ───────────────────────────────────────────────────────── │
│ 1. Agy 49d 49d 59d 214.2K step │
│ 2. command-code 1d 9d 49d 11.2K msg │
│ 3. Droid 1d 1d 2d 14 msg │
│ 4. Muse 1d 1d 1d 1 prompt │
╰────────────────────────────────────────────────────────────╯
>> Agy is on fire with a 49-day streak!
Recorded cost: $119.52 from 4 tools. Estimated: ~$20697.60 across 5 tools (~ prefix).
Seeing the raw numbers changes how you work. You immediately notice when a harness gets stuck in a repetitive prompt loop, or when a model switch cuts daily expenditure by two orders of magnitude. For single-tool deep dives, Thermal renders eight-week activity heatmaps directly in the terminal, drawing inspiration from contribution graphs:
$ thermal --tool opencode --weeks 8
╭─ OpenCode Activity ───────────────────────────────────── ~/.local/share/opencode/opencode.db ─╮
│ 10.6B tokens / 8 weeks · 30 active days | 1 day streak | 11 best | 10.6B all-time │
│ │
│ Aug Sep │
│ ■■■■■■■■ │
│ Mon ■■■■■■■■ │
│ ■■■■■■■■ │
│ Wed ■■■■■■■■ │
│ ■■■■■■■ │
│ Fri ■■■■■■■ │
│ ■■■■■■■ │
│ Less ■■■■■ More │
│ ──────────────────────────────────────────────────────────────────────────────────────────── │
│ $62.91 spent · 98% cache · 256 sessions │
│ · agents build: 95 reviewer: 59 general: 58 │
╰───────────────────────────────────────────────────────────────────────────────────────────────╯
As I wrote earlier in my essay on why coding agents burn tokens in the dark, visibility is the prerequisite for discipline. You cannot optimize what you cannot measure.
Thermal is distributed as an open-source Go binary:
go install github.com/jadmadi/thermal/cmd/thermal@latest
Skill Cabinet Go: Ordering the Drawers
Once you can see token consumption, the next bottleneck becomes context pollution.
Every agent tool now supports the concept of skills: specialized Markdown
instruction files with YAML frontmatter that tell the model how to perform
specific tasks, such as auditing database migrations or debugging memory
leaks. But every tool also insists on its own directory convention.
Cursor uses .cursor/, Claude Code uses .claude/, Codex uses .codex/,
and cross-agent tools look in .agents/.
When you maintain dozens of projects and hundreds of skills, this setup devolves into chaos. You either copy skill folders by hand across fifty drawers, creating stale forks that drift over time, or you symlink everything into a giant global bucket. As I documented in The Skill Bloat Dilemma, loading several hundred skills indiscriminately into an agent system prompt causes silent alphabetical truncation past fifty tools and consumes fifty thousand tokens before you have typed your first instruction.
To solve this, I built skill-cabinet-go, a high-performance Go companion
to the open-source subsy/skill-cabinet standard.
Skill Cabinet treats your machine as a house with distinct drawers. It scans every tool drawer, detects physical files versus symlinks, flags broken links, audits executable script permissions, and deduplicates identical skills.
Running a census across my workstation reveals the scale of the problem:
$ skill-cabinet-go count
skill-cabinet-go 0.3.0 (commit c6260ed, built 2026-09-02T17:33:20Z)
in house 10574 · physical 667 · references 9599 · broken 308 · 90 copies · 5.5 MB
More than ten thousand skill cards live across fifty-eight different drawers on my workstation. Without a tool to inspect them, three hundred broken symlinks would be quietly failing inside agent tool registries without warning.
With skill-cabinet-go, I can inspect individual drawers, verify their risk
ratings, and check modification dates at a glance:
$ skill-cabinet-go ls --scope codex
drawer name storage risk copies modified
------ ------------------------- --------- -------- ------ ----------
.codex adapt physical none 1 2026-09-21
.codex agents-sdk physical none 3 2026-09-14
.codex animate physical none 1 2026-09-21
.codex audit physical none 1 2026-09-21
.codex bolder physical none 1 2026-09-21
.codex brief physical none 1 2026-09-21
.codex clarify physical none 1 2026-09-21
.codex cloudflare physical none 1 2026-09-14
.codex cloudflare-email-service physical none 3 2026-09-14
.codex cloudflare-one reference none - 2026-08-15
.codex cloudflare-one-migrations reference none - 2026-08-15
.codex colorize physical none 1 2026-09-21
.codex craft physical none 1 2026-09-21
.codex critique physical none 1 2026-09-21
.codex delight physical none 1 2026-09-21
.codex distill physical none 1 2026-09-21
.codex durable-objects physical none 3 2026-09-14
.codex extract physical none 1 2026-09-21
Instead of spraying skills everywhere, I keep a single canonical source of
truth and use skill-cabinet-go link and dedupe to synchronize
references. The binary also exposes a native MCP server (skill-cabinet-go mcp),
allowing agents to search and inspect skills dynamically rather than loading
all six hundred physical cards into the static system prompt.
The code and specification live at github.com/jadmadi/skill-cabinet-go.
The Judgment Layer: Agent Skills That Guard the Work
Having meters and tidy drawers is still just mechanical hygiene. The real difference in day-to-day coding comes from the procedural judgment you hand to the model.
In github.com/jadmadi/skills, I maintain the core operational playbooks
that my agents load during work. As explored in
Demystifying Agent Skills, an effective
skill is not a polite prompt asking the model to write clean code. It is an
opinionated execution harness that restricts choices and enforces empirical
verification.
Two skills in particular anchor almost every session:
-
Pre-flight Check (
pre-flight-check): A pre-commit and pre-deploy checklist for backend and full-stack modifications. It forces the agent to verify API response envelopes, check route ordering so static routes are not swallowed by dynamic wildcards, verify database column schemas before issuing writes, and refuse to commit code without verifying local exit code zero. -
Doubt-Driven Development & Context Engineering: Skills that require the agent to cross-examine its own assumptions before touching shared interfaces. When an agent encounters an unexpected error, instead of guessing random workarounds or modifying configuration files at random, the skill directs it to check persistent project memory and isolate root causes.
When you pair procedural skills with native MCP tools, the agent stops acting like an overeager junior intern writing speculative code. It behaves like a disciplined pair programmer who reads the manual, checks the database schema, and writes verified tests.
Where the Scaffolding Led: Mahak and Sila
Building this local harness changed how I think about larger AI systems. Once you start measuring things locally with real numbers, you realize how much of the wider industry still relies on vibes and marketing claims.
That realization led directly to two larger projects.
Mahak: Measuring Arabic Fluency Empirically
When evaluating models for non-English languages, especially Arabic, public benchmarks are notoriously misleading. Synthetic multiple-choice tests rarely reflect whether a model can draft a legally sound contract in contemporary Arabic or handle nuanced technical discourse without awkward literal translations.
Applying the same empirical philosophy behind Thermal, I built Mahak (مَحَكّ), an open community benchmark ranking frontier and open-weights models on native Arabic fluency across practical domains.
On the live Results Matrix today, Mahak indexes 65 frontier and open-weights models across 25 rigorous tasks spanning legal contracts, creative writing, customer support, business communication, and instruction following. Running the evaluation CLI on this very machine:
mahak start --model claude-3-7-sonnet --domains all
collects blind model outputs that enter community evaluation and Elo scoring. No marketing spin, no cherry-picked prompts. Just transparent, verifiable results.
Sila: Multi-Agent Continuity and Goal Governance
The second evolution addressed the biggest limitation of all: agent amnesia.
When you switch between different coding sessions or collaborate across multiple specialized agents, knowledge evaporates. One agent solves an obscure SQLite concurrency issue on Tuesday; by Thursday, a different agent hits the exact same error and burns thirty minutes rediscovering the fix.
To solve this cross-harness continuity gap, I created Sila (صِلَة). Sila is a multi-agent governance and persistent memory substrate built on local SQLite stores. It continuously indexes session messages across tools like OpenCode, Devin, Codex, and Claude Code, distills reusable lessons, and coordinates roadmaps through an explicit goal system.
Before starting work on this very essay, I checked the local goal status:
$ sila goals
---------------------------------------------------------------
🟢 RECENTLY IMPLEMENTED & VERIFIED (3):
---------------------------------------------------------------
✔ 🚢 ⚡ seo-audit-fixes [commit e2a5954] (11 tasks)
✔ 🚢 ⚡ sila-onboarding [commit a846e1e] (4 tasks)
✔ 🚢 ⚡ webmcp-projects [commit 99580a3] (5 tasks)
🟡 GOALS READY TO CLAIM & EXECUTE (1):
---------------------------------------------------------------
▶ mcp-stack-post 🚢 🗺️ [Roadmap] ⚡ [Independent]
Title: MCP stack post : how I run my own agent tools
Status: ready to claim; awaiting implementation commit
Claim: sila goals claim mcp-stack-post --agent=<name>
Plan: [░░░░░░░░] 0/9 tasks (0%)
---------------------------------------------------------------
The agent claims a goal, locks its boundaries, completes the planned steps, and verifies the done conditions before recording a verifiable handoff. Session context does not get lost in terminal scrollback; it enters a permanent, queryable local knowledge base.
Sila is currently in active private development as I refine its multi-tool ingestion pipelines and MCP tool surfaces, and I will share more about its architecture when it is ready for public preview.
The Local-First Philosophy
If there is a unifying thread across these tools, it is a stubborn commitment to local-first software.
We do not need to ship every keystroke and prompt to opaque cloud services just to benefit from intelligent coding assistants. The most reliable, responsive, and private developer experience happens right here on your own machine. When you give agents fast, compiled tools written in Go, back them with lightweight SQLite databases, and govern them with clear procedural skills, you reclaim control over your environment.
You stop wondering why your API bill doubled. You stop dreading context exhaustion. And you start treating coding agents for what they truly can be: sharp, dependable partners in the craft of building software.
The Stack at a Glance
For those who want to explore or run these tools locally:
-
Thermal: Local-first agent FinOps and token tracking.
Repository: github.com/jadmadi/thermal
Install:go install github.com/jadmadi/thermal/cmd/thermal@latest
Project Overview: /project/thermal/ -
Skill Cabinet Go: Census, audit, and deduplication for agent skills.
Repository: github.com/jadmadi/skill-cabinet-go
Upstream Standard: subsy/skill-cabinet -
Agent Skills: Curated procedural skills and verification harnesses.
Repository: github.com/jadmadi/skills -
Mahak Benchmark: Open Arabic fluency benchmark and evaluation CLI.
Live Matrix: mahak.waqf.dev/en/matrix/ -
Profile & Catalog: Overview of open-source tooling and standards.
GitHub: github.com/jadmadi
Projects Catalog: /projects/
Recommended
Continue reading
Selected from shared topics, related tags, and the recent archive.