My MCP Stack: How I Run My Own Agent Tools

By Jad • • 11 min read
My MCP Stack: How I Run My Own Agent Tools

A few months ago, I caught myself looking at a completed pull request with a lingering sense of unease.

The diff itself was clean. Twelve lines of Go in an HTTP router, a straightforward unit test, and a green badge on the CI runner. If you inspected the git commit log, the entire change looked like a ten-minute task.

What the git log failed to record was the invisible churn that produced it.

Behind those twelve lines was an autonomous agent session that ran for nearly forty minutes. It had traversed eighteen files, triggered thirty-two model turns, burned through four hundred thousand tokens, and silently stumbled down two architectural dead ends before finding the right file. I was paying the API invoice for an entire engineering afternoon, yet my tools treated the work as if it had cost nothing more than a few keystrokes.

We are living through a curious transition in software engineering. We have welcomed autonomous coding agents into our daily workflow, yet we still manage them with nineteenth-century habits and twentieth-century terminal tooling. We know the exact memory footprint of our compiled binaries, but we have zero visibility into the token burn of our daily coding sessions. We collect procedural agent skills like browser bookmarks, scattering them across dozens of vendor-specific dotfile directories until our context windows choke on alphabetical truncation.

Over the past year, I stopped treating agents as black-box appliances and started building the missing scaffolding. The goal was never to wrap them in complicated enterprise orchestration platforms. Instead, I wanted small, local-first utilities built in Go that run on my own workstation, read ordinary SQLite databases, speak standard protocols like the Model Context Protocol (MCP), and answer three basic questions: What did this session cost, what skills are my agents actually carrying, and how do we keep their work grounded in reality?

Here is the stack I run every day.

Thermal: Financial Vision from Read-Only SQLite

The first missing piece was accounting. When an agent runs in your terminal, it writes session telemetry somewhere on your local disk. OpenCode, Devin, Codex, MiMoCode, and Claude each maintain local stores. Some use JSON lines, others use SQLite, and a few store raw session state in proprietary caches.

Rather than running an intercepting proxy daemon that injects network latency or wrapping agent binaries in fragile shell shims, I built Thermal.

Thermal is an open-source terminal utility written in Go. Its operating philosophy is strictly read-only: it reaches directly into the local databases already left behind by your coding tools on disk, parses token usage and model identifiers, and calculates running streaks, daily volume, and real financial spend. Because it reads local files directly with SQLite memory-mapped I/O, generating an entire cross-tool census takes less than twenty milliseconds.

When I run thermal in my workspace today, it does not give me vague estimates or cloud dashboard spinners. It prints an immediate, concrete leaderboard:

  THERMAL : Don't break the streak.

  ╭─ Token Warriors ────────────────────────────────────── tokens & spend ─╮
  │    #  Tool              Strk    Best    Days       Tokens         Cost │
  │  ───────────────────────────────────────────────────────────────────── │
  │   1.  Codex               3d      7d     49d     1.0B tok     ~$538.30 │
  │   2.  Devin               2d     33d     45d    19.2B tok   ~$20134.07 │
  │   3.  OpenCode            1d     11d     30d    10.6B tok       $62.91 │
  │   4.  MiMoCode            1d      7d     16d     1.2B tok       $56.14 │
  │   5.  ZCode               1d      6d     11d   871.0M tok      ~$25.01 │
  │   6.  DeepSeek (DSH)      1d      1d      5d    18.0M tok       ~$0.07 │
  │   7.  Grok                1d      1d      1d   692.7K tok        $0.44 │
  │   8.  codewhale           1d      1d      3d   110.1K tok        $0.02 │
  │   9.  Claude              1d      1d      3d    55.5K tok       ~$0.16 │
  ╰────────────────────────────────────────────────────────────────────────╯

  ╭─ Activity Hunters ────────────────────── messages & steps ─╮
  │    #  Tool              Strk    Best    Days      Activity │
  │  ───────────────────────────────────────────────────────── │
  │   1.  Agy                49d     49d     59d   214.2K step │
  │   2.  command-code        1d      9d     49d     11.2K msg │
  │   3.  Droid               1d      1d      2d        14 msg │
  │   4.  Muse                1d      1d      1d      1 prompt │
  ╰────────────────────────────────────────────────────────────╯

  >> Agy is on fire with a 49-day streak!

  Recorded cost: $119.52 from 4 tools. Estimated: ~$20697.60 across 5 tools (~ prefix).

Seeing the raw numbers changes how you work. You immediately notice when a harness gets stuck in a repetitive prompt loop, or when a model switch cuts daily expenditure by two orders of magnitude. For single-tool deep dives, Thermal renders eight-week activity heatmaps directly in the terminal, drawing inspiration from contribution graphs:

$ thermal --tool opencode --weeks 8

  ╭─ OpenCode Activity ───────────────────────────────────── ~/.local/share/opencode/opencode.db ─╮
  │  10.6B tokens / 8 weeks   ·   30 active days  |  1 day streak  |  11 best  |  10.6B all-time  │
  │                                                                                               │
  │       Aug Sep                                                                                 │
  │       ■■■■■■■■                                                                                │
  │   Mon ■■■■■■■■                                                                                │
  │       ■■■■■■■■                                                                                │
  │   Wed ■■■■■■■■                                                                                │
  │       ■■■■■■■                                                                                 │
  │   Fri ■■■■■■■                                                                                 │
  │       ■■■■■■■                                                                                 │
  │       Less ■■■■■ More                                                                         │
  │  ──────────────────────────────────────────────────────────────────────────────────────────── │
  │  $62.91 spent  ·  98% cache  ·  256 sessions                                                  │
  │  · agents  build: 95  reviewer: 59  general: 58                                               │
  ╰───────────────────────────────────────────────────────────────────────────────────────────────╯

As I wrote earlier in my essay on why coding agents burn tokens in the dark, visibility is the prerequisite for discipline. You cannot optimize what you cannot measure.

Thermal is distributed as an open-source Go binary:

go install github.com/jadmadi/thermal/cmd/thermal@latest

Skill Cabinet Go: Ordering the Drawers

Once you can see token consumption, the next bottleneck becomes context pollution.

Every agent tool now supports the concept of skills: specialized Markdown instruction files with YAML frontmatter that tell the model how to perform specific tasks, such as auditing database migrations or debugging memory leaks. But every tool also insists on its own directory convention. Cursor uses .cursor/, Claude Code uses .claude/, Codex uses .codex/, and cross-agent tools look in .agents/.

When you maintain dozens of projects and hundreds of skills, this setup devolves into chaos. You either copy skill folders by hand across fifty drawers, creating stale forks that drift over time, or you symlink everything into a giant global bucket. As I documented in The Skill Bloat Dilemma, loading several hundred skills indiscriminately into an agent system prompt causes silent alphabetical truncation past fifty tools and consumes fifty thousand tokens before you have typed your first instruction.

To solve this, I built skill-cabinet-go, a high-performance Go companion to the open-source subsy/skill-cabinet standard.

Skill Cabinet treats your machine as a house with distinct drawers. It scans every tool drawer, detects physical files versus symlinks, flags broken links, audits executable script permissions, and deduplicates identical skills.

Running a census across my workstation reveals the scale of the problem:

$ skill-cabinet-go count
skill-cabinet-go 0.3.0 (commit c6260ed, built 2026-09-02T17:33:20Z)
in house 10574 · physical 667 · references 9599 · broken 308 · 90 copies · 5.5 MB

More than ten thousand skill cards live across fifty-eight different drawers on my workstation. Without a tool to inspect them, three hundred broken symlinks would be quietly failing inside agent tool registries without warning.

With skill-cabinet-go, I can inspect individual drawers, verify their risk ratings, and check modification dates at a glance:

$ skill-cabinet-go ls --scope codex
drawer  name                       storage    risk      copies  modified
------  -------------------------  ---------  --------  ------  ----------
.codex  adapt                      physical   none      1       2026-09-21
.codex  agents-sdk                 physical   none      3       2026-09-14
.codex  animate                    physical   none      1       2026-09-21
.codex  audit                      physical   none      1       2026-09-21
.codex  bolder                     physical   none      1       2026-09-21
.codex  brief                      physical   none      1       2026-09-21
.codex  clarify                    physical   none      1       2026-09-21
.codex  cloudflare                 physical   none      1       2026-09-14
.codex  cloudflare-email-service   physical   none      3       2026-09-14
.codex  cloudflare-one             reference  none      -       2026-08-15
.codex  cloudflare-one-migrations  reference  none      -       2026-08-15
.codex  colorize                   physical   none      1       2026-09-21
.codex  craft                      physical   none      1       2026-09-21
.codex  critique                   physical   none      1       2026-09-21
.codex  delight                    physical   none      1       2026-09-21
.codex  distill                    physical   none      1       2026-09-21
.codex  durable-objects            physical   none      3       2026-09-14
.codex  extract                    physical   none      1       2026-09-21

Instead of spraying skills everywhere, I keep a single canonical source of truth and use skill-cabinet-go link and dedupe to synchronize references. The binary also exposes a native MCP server (skill-cabinet-go mcp), allowing agents to search and inspect skills dynamically rather than loading all six hundred physical cards into the static system prompt.

The code and specification live at github.com/jadmadi/skill-cabinet-go.

The Judgment Layer: Agent Skills That Guard the Work

Having meters and tidy drawers is still just mechanical hygiene. The real difference in day-to-day coding comes from the procedural judgment you hand to the model.

In github.com/jadmadi/skills, I maintain the core operational playbooks that my agents load during work. As explored in Demystifying Agent Skills, an effective skill is not a polite prompt asking the model to write clean code. It is an opinionated execution harness that restricts choices and enforces empirical verification.

Two skills in particular anchor almost every session:

  1. Pre-flight Check (pre-flight-check): A pre-commit and pre-deploy checklist for backend and full-stack modifications. It forces the agent to verify API response envelopes, check route ordering so static routes are not swallowed by dynamic wildcards, verify database column schemas before issuing writes, and refuse to commit code without verifying local exit code zero.

  2. Doubt-Driven Development & Context Engineering: Skills that require the agent to cross-examine its own assumptions before touching shared interfaces. When an agent encounters an unexpected error, instead of guessing random workarounds or modifying configuration files at random, the skill directs it to check persistent project memory and isolate root causes.

When you pair procedural skills with native MCP tools, the agent stops acting like an overeager junior intern writing speculative code. It behaves like a disciplined pair programmer who reads the manual, checks the database schema, and writes verified tests.

Where the Scaffolding Led: Mahak and Sila

Building this local harness changed how I think about larger AI systems. Once you start measuring things locally with real numbers, you realize how much of the wider industry still relies on vibes and marketing claims.

That realization led directly to two larger projects.

Mahak: Measuring Arabic Fluency Empirically

When evaluating models for non-English languages, especially Arabic, public benchmarks are notoriously misleading. Synthetic multiple-choice tests rarely reflect whether a model can draft a legally sound contract in contemporary Arabic or handle nuanced technical discourse without awkward literal translations.

Applying the same empirical philosophy behind Thermal, I built Mahak (مَحَكّ), an open community benchmark ranking frontier and open-weights models on native Arabic fluency across practical domains.

On the live Results Matrix today, Mahak indexes 65 frontier and open-weights models across 25 rigorous tasks spanning legal contracts, creative writing, customer support, business communication, and instruction following. Running the evaluation CLI on this very machine:

mahak start --model claude-3-7-sonnet --domains all

collects blind model outputs that enter community evaluation and Elo scoring. No marketing spin, no cherry-picked prompts. Just transparent, verifiable results.

Sila: Multi-Agent Continuity and Goal Governance

The second evolution addressed the biggest limitation of all: agent amnesia.

When you switch between different coding sessions or collaborate across multiple specialized agents, knowledge evaporates. One agent solves an obscure SQLite concurrency issue on Tuesday; by Thursday, a different agent hits the exact same error and burns thirty minutes rediscovering the fix.

To solve this cross-harness continuity gap, I created Sila (صِلَة). Sila is a multi-agent governance and persistent memory substrate built on local SQLite stores. It continuously indexes session messages across tools like OpenCode, Devin, Codex, and Claude Code, distills reusable lessons, and coordinates roadmaps through an explicit goal system.

Before starting work on this very essay, I checked the local goal status:

$ sila goals
---------------------------------------------------------------
🟢 RECENTLY IMPLEMENTED & VERIFIED (3):
---------------------------------------------------------------
  ✔ 🚢 ⚡ seo-audit-fixes                            [commit e2a5954] (11 tasks)
  ✔ 🚢 ⚡ sila-onboarding                            [commit a846e1e] (4 tasks)
  ✔ 🚢 ⚡ webmcp-projects                            [commit 99580a3] (5 tasks)

🟡 GOALS READY TO CLAIM & EXECUTE (1):
---------------------------------------------------------------
  ▶ mcp-stack-post 🚢 🗺️ [Roadmap] ⚡ [Independent]
    Title:  MCP stack post : how I run my own agent tools
    Status: ready to claim; awaiting implementation commit
    Claim:  sila goals claim mcp-stack-post --agent=<name>
    Plan:   [░░░░░░░░] 0/9 tasks (0%)
---------------------------------------------------------------

The agent claims a goal, locks its boundaries, completes the planned steps, and verifies the done conditions before recording a verifiable handoff. Session context does not get lost in terminal scrollback; it enters a permanent, queryable local knowledge base.

Sila is currently in active private development as I refine its multi-tool ingestion pipelines and MCP tool surfaces, and I will share more about its architecture when it is ready for public preview.

The Local-First Philosophy

If there is a unifying thread across these tools, it is a stubborn commitment to local-first software.

We do not need to ship every keystroke and prompt to opaque cloud services just to benefit from intelligent coding assistants. The most reliable, responsive, and private developer experience happens right here on your own machine. When you give agents fast, compiled tools written in Go, back them with lightweight SQLite databases, and govern them with clear procedural skills, you reclaim control over your environment.

You stop wondering why your API bill doubled. You stop dreading context exhaustion. And you start treating coding agents for what they truly can be: sharp, dependable partners in the craft of building software.


The Stack at a Glance

For those who want to explore or run these tools locally:

Recommended

Selected from shared topics, related tags, and the recent archive.