---
title: 'My MCP Stack: How I Run My Own Agent Tools'
description: >-
  A peek into my daily local-first agent workflow across Thermal, Skill Cabinet
  Go, and curated agent skills. How tracking token burn and organizing skill
  drawers grew into Mahak and Sila.
summary: >-
  Autonomous coding agents need local discipline: honest token accounting,
  governed skill drawers, and repeatable quality gates. Here is how I run
  Thermal, Skill Cabinet Go, and agent skills on my own machine.
date: 2026-09-30T00:00:00.000Z
heroImage:
  src: /_astro/hero.D8TWuCo5.webp
  width: 1360
  height: 765
  format: webp
tags:
  - ai-agents
  - local-first
  - go
  - sqlite
  - mcp
categories:
  - Technology Bites
published: true
featured: false
draft: false
author: Jad
href: /posts/my-mcp-stack/
slug: my-mcp-stack
---
A few months ago, I caught myself looking at a completed pull request
with a lingering sense of unease.

The diff itself was clean. Twelve lines of Go in an HTTP router, a
straightforward unit test, and a green badge on the CI runner. If you
inspected the git commit log, the entire change looked like a ten-minute
task.

What the git log failed to record was the invisible churn that produced
it.

Behind those twelve lines was an autonomous agent session that ran for
nearly forty minutes. It had traversed eighteen files, triggered thirty-two
model turns, burned through four hundred thousand tokens, and silently
stumbled down two architectural dead ends before finding the right file.
I was paying the API invoice for an entire engineering afternoon, yet my
tools treated the work as if it had cost nothing more than a few
keystrokes.

We are living through a curious transition in software engineering. We have
welcomed autonomous coding agents into our daily workflow, yet we still
manage them with nineteenth-century habits and twentieth-century terminal
tooling. We know the exact memory footprint of our compiled binaries, but we
have zero visibility into the token burn of our daily coding sessions. We
collect procedural agent skills like browser bookmarks, scattering them
across dozens of vendor-specific dotfile directories until our context
windows choke on alphabetical truncation.

Over the past year, I stopped treating agents as black-box appliances and
started building the missing scaffolding. The goal was never to wrap them
in complicated enterprise orchestration platforms. Instead, I wanted small,
local-first utilities built in Go that run on my own workstation, read
ordinary SQLite databases, speak standard protocols like the Model Context
Protocol (MCP), and answer three basic questions: What did this session
cost, what skills are my agents actually carrying, and how do we keep their
work grounded in reality?

Here is the stack I run every day.

## Thermal: Financial Vision from Read-Only SQLite

The first missing piece was accounting. When an agent runs in your terminal,
it writes session telemetry somewhere on your local disk. OpenCode, Devin,
Codex, MiMoCode, and Claude each maintain local stores. Some use JSON lines,
others use SQLite, and a few store raw session state in proprietary caches.

Rather than running an intercepting proxy daemon that injects network latency
or wrapping agent binaries in fragile shell shims, I built [Thermal](/project/thermal/).

Thermal is an open-source terminal utility written in Go. Its operating
philosophy is strictly read-only: it reaches directly into the local
databases already left behind by your coding tools on disk, parses token
usage and model identifiers, and calculates running streaks, daily volume,
and real financial spend. Because it reads local files directly with SQLite
memory-mapped I/O, generating an entire cross-tool census takes less than
twenty milliseconds.

When I run `thermal` in my workspace today, it does not give me vague
estimates or cloud dashboard spinners. It prints an immediate, concrete
leaderboard:

```text
  THERMAL : Don't break the streak.

  ╭─ Token Warriors ────────────────────────────────────── tokens & spend ─╮
  │    #  Tool              Strk    Best    Days       Tokens         Cost │
  │  ───────────────────────────────────────────────────────────────────── │
  │   1.  Codex               3d      7d     49d     1.0B tok     ~$538.30 │
  │   2.  Devin               2d     33d     45d    19.2B tok   ~$20134.07 │
  │   3.  OpenCode            1d     11d     30d    10.6B tok       $62.91 │
  │   4.  MiMoCode            1d      7d     16d     1.2B tok       $56.14 │
  │   5.  ZCode               1d      6d     11d   871.0M tok      ~$25.01 │
  │   6.  DeepSeek (DSH)      1d      1d      5d    18.0M tok       ~$0.07 │
  │   7.  Grok                1d      1d      1d   692.7K tok        $0.44 │
  │   8.  codewhale           1d      1d      3d   110.1K tok        $0.02 │
  │   9.  Claude              1d      1d      3d    55.5K tok       ~$0.16 │
  ╰────────────────────────────────────────────────────────────────────────╯

  ╭─ Activity Hunters ────────────────────── messages & steps ─╮
  │    #  Tool              Strk    Best    Days      Activity │
  │  ───────────────────────────────────────────────────────── │
  │   1.  Agy                49d     49d     59d   214.2K step │
  │   2.  command-code        1d      9d     49d     11.2K msg │
  │   3.  Droid               1d      1d      2d        14 msg │
  │   4.  Muse                1d      1d      1d      1 prompt │
  ╰────────────────────────────────────────────────────────────╯

  >> Agy is on fire with a 49-day streak!

  Recorded cost: $119.52 from 4 tools. Estimated: ~$20697.60 across 5 tools (~ prefix).
```

Seeing the raw numbers changes how you work. You immediately notice when a
harness gets stuck in a repetitive prompt loop, or when a model switch cuts
daily expenditure by two orders of magnitude. For single-tool deep dives,
Thermal renders eight-week activity heatmaps directly in the terminal,
drawing inspiration from contribution graphs:

```text
$ thermal --tool opencode --weeks 8

  ╭─ OpenCode Activity ───────────────────────────────────── ~/.local/share/opencode/opencode.db ─╮
  │  10.6B tokens / 8 weeks   ·   30 active days  |  1 day streak  |  11 best  |  10.6B all-time  │
  │                                                                                               │
  │       Aug Sep                                                                                 │
  │       ■■■■■■■■                                                                                │
  │   Mon ■■■■■■■■                                                                                │
  │       ■■■■■■■■                                                                                │
  │   Wed ■■■■■■■■                                                                                │
  │       ■■■■■■■                                                                                 │
  │   Fri ■■■■■■■                                                                                 │
  │       ■■■■■■■                                                                                 │
  │       Less ■■■■■ More                                                                         │
  │  ──────────────────────────────────────────────────────────────────────────────────────────── │
  │  $62.91 spent  ·  98% cache  ·  256 sessions                                                  │
  │  · agents  build: 95  reviewer: 59  general: 58                                               │
  ╰───────────────────────────────────────────────────────────────────────────────────────────────╯
```

As I wrote earlier in my essay on
[why coding agents burn tokens in the dark](/posts/why-your-coding-agents-burn-tokens-in-the-dark/),
visibility is the prerequisite for discipline. You cannot optimize what you
cannot measure.

Thermal is distributed as an open-source Go binary:

```bash
go install github.com/jadmadi/thermal/cmd/thermal@latest
```

## Skill Cabinet Go: Ordering the Drawers

Once you can see token consumption, the next bottleneck becomes context
pollution.

Every agent tool now supports the concept of skills: specialized Markdown
instruction files with YAML frontmatter that tell the model how to perform
specific tasks, such as auditing database migrations or debugging memory
leaks. But every tool also insists on its own directory convention.
Cursor uses `.cursor/`, Claude Code uses `.claude/`, Codex uses `.codex/`,
and cross-agent tools look in `.agents/`.

When you maintain dozens of projects and hundreds of skills, this setup
devolves into chaos. You either copy skill folders by hand across fifty
drawers, creating stale forks that drift over time, or you symlink
everything into a giant global bucket. As I documented in
[The Skill Bloat Dilemma](/posts/the-skill-bloat-dilemma/), loading several
hundred skills indiscriminately into an agent system prompt causes silent
alphabetical truncation past fifty tools and consumes fifty thousand tokens
before you have typed your first instruction.

To solve this, I built `skill-cabinet-go`, a high-performance Go companion
to the open-source `subsy/skill-cabinet` standard.

Skill Cabinet treats your machine as a house with distinct drawers. It
scans every tool drawer, detects physical files versus symlinks, flags broken
links, audits executable script permissions, and deduplicates identical
skills.

Running a census across my workstation reveals the scale of the problem:

```text
$ skill-cabinet-go count
skill-cabinet-go 0.3.0 (commit c6260ed, built 2026-09-02T17:33:20Z)
in house 10574 · physical 667 · references 9599 · broken 308 · 90 copies · 5.5 MB
```

More than ten thousand skill cards live across fifty-eight different
drawers on my workstation. Without a tool to inspect them, three hundred
broken symlinks would be quietly failing inside agent tool registries
without warning.

With `skill-cabinet-go`, I can inspect individual drawers, verify their risk
ratings, and check modification dates at a glance:

```text
$ skill-cabinet-go ls --scope codex
drawer  name                       storage    risk      copies  modified
------  -------------------------  ---------  --------  ------  ----------
.codex  adapt                      physical   none      1       2026-09-21
.codex  agents-sdk                 physical   none      3       2026-09-14
.codex  animate                    physical   none      1       2026-09-21
.codex  audit                      physical   none      1       2026-09-21
.codex  bolder                     physical   none      1       2026-09-21
.codex  brief                      physical   none      1       2026-09-21
.codex  clarify                    physical   none      1       2026-09-21
.codex  cloudflare                 physical   none      1       2026-09-14
.codex  cloudflare-email-service   physical   none      3       2026-09-14
.codex  cloudflare-one             reference  none      -       2026-08-15
.codex  cloudflare-one-migrations  reference  none      -       2026-08-15
.codex  colorize                   physical   none      1       2026-09-21
.codex  craft                      physical   none      1       2026-09-21
.codex  critique                   physical   none      1       2026-09-21
.codex  delight                    physical   none      1       2026-09-21
.codex  distill                    physical   none      1       2026-09-21
.codex  durable-objects            physical   none      3       2026-09-14
.codex  extract                    physical   none      1       2026-09-21
```

Instead of spraying skills everywhere, I keep a single canonical source of
truth and use `skill-cabinet-go link` and `dedupe` to synchronize
references. The binary also exposes a native MCP server (`skill-cabinet-go mcp`),
allowing agents to search and inspect skills dynamically rather than loading
all six hundred physical cards into the static system prompt.

The code and specification live at `github.com/jadmadi/skill-cabinet-go`.

## The Judgment Layer: Agent Skills That Guard the Work

Having meters and tidy drawers is still just mechanical hygiene. The real
difference in day-to-day coding comes from the procedural judgment you hand
to the model.

In `github.com/jadmadi/skills`, I maintain the core operational playbooks
that my agents load during work. As explored in
[Demystifying Agent Skills](/posts/demystifying-agent-skills/), an effective
skill is not a polite prompt asking the model to write clean code. It is an
opinionated execution harness that restricts choices and enforces empirical
verification.

Two skills in particular anchor almost every session:

1. **Pre-flight Check (`pre-flight-check`):**
   A pre-commit and pre-deploy checklist for backend and full-stack
   modifications. It forces the agent to verify API response envelopes,
   check route ordering so static routes are not swallowed by dynamic
   wildcards, verify database column schemas before issuing writes, and
   refuse to commit code without verifying local exit code zero.

2. **Doubt-Driven Development & Context Engineering:**
   Skills that require the agent to cross-examine its own assumptions before
   touching shared interfaces. When an agent encounters an unexpected error,
   instead of guessing random workarounds or modifying configuration files
   at random, the skill directs it to check persistent project memory and
   isolate root causes.

When you pair procedural skills with native MCP tools, the agent stops acting
like an overeager junior intern writing speculative code. It behaves like a
disciplined pair programmer who reads the manual, checks the database schema,
and writes verified tests.

## Where the Scaffolding Led: Mahak and Sila

Building this local harness changed how I think about larger AI systems.
Once you start measuring things locally with real numbers, you realize how
much of the wider industry still relies on vibes and marketing claims.

That realization led directly to two larger projects.

### Mahak: Measuring Arabic Fluency Empirically

When evaluating models for non-English languages, especially Arabic, public
benchmarks are notoriously misleading. Synthetic multiple-choice tests
rarely reflect whether a model can draft a legally sound contract in
contemporary Arabic or handle nuanced technical discourse without awkward
literal translations.

Applying the same empirical philosophy behind Thermal, I built
[Mahak (مَحَكّ)](https://mahak.waqf.dev), an open community benchmark
ranking frontier and open-weights models on native Arabic fluency across
practical domains.

On the live [Results Matrix](https://mahak.waqf.dev/en/matrix/) today, Mahak
indexes 65 frontier and open-weights models across 25 rigorous tasks spanning
legal contracts, creative writing, customer support, business
communication, and instruction following. Running the evaluation CLI on this
very machine:

```bash
mahak start --model claude-3-7-sonnet --domains all
```

collects blind model outputs that enter community evaluation and Elo scoring.
No marketing spin, no cherry-picked prompts. Just transparent, verifiable
results.

### Sila: Multi-Agent Continuity and Goal Governance

The second evolution addressed the biggest limitation of all: agent amnesia.

When you switch between different coding sessions or collaborate across
multiple specialized agents, knowledge evaporates. One agent solves an
obscure SQLite concurrency issue on Tuesday; by Thursday, a different agent
hits the exact same error and burns thirty minutes rediscovering the fix.

To solve this cross-harness continuity gap, I created Sila (صِلَة). Sila is a
multi-agent governance and persistent memory substrate built on local
SQLite stores. It continuously indexes session messages across tools like
OpenCode, Devin, Codex, and Claude Code, distills reusable lessons, and
coordinates roadmaps through an explicit goal system.

Before starting work on this very essay, I checked the local goal status:

```text
$ sila goals
---------------------------------------------------------------
🟢 RECENTLY IMPLEMENTED & VERIFIED (3):
---------------------------------------------------------------
  ✔ 🚢 ⚡ seo-audit-fixes                            [commit e2a5954] (11 tasks)
  ✔ 🚢 ⚡ sila-onboarding                            [commit a846e1e] (4 tasks)
  ✔ 🚢 ⚡ webmcp-projects                            [commit 99580a3] (5 tasks)

🟡 GOALS READY TO CLAIM & EXECUTE (1):
---------------------------------------------------------------
  ▶ mcp-stack-post 🚢 🗺️ [Roadmap] ⚡ [Independent]
    Title:  MCP stack post : how I run my own agent tools
    Status: ready to claim; awaiting implementation commit
    Claim:  sila goals claim mcp-stack-post --agent=<name>
    Plan:   [░░░░░░░░] 0/9 tasks (0%)
---------------------------------------------------------------
```

The agent claims a goal, locks its boundaries, completes the planned steps,
and verifies the done conditions before recording a verifiable handoff.
Session context does not get lost in terminal scrollback; it enters a
permanent, queryable local knowledge base.

Sila is currently in active private development as I refine its multi-tool
ingestion pipelines and MCP tool surfaces, and I will share more about its
architecture when it is ready for public preview.

## The Local-First Philosophy

If there is a unifying thread across these tools, it is a stubborn
commitment to local-first software.

We do not need to ship every keystroke and prompt to opaque cloud services
just to benefit from intelligent coding assistants. The most reliable,
responsive, and private developer experience happens right here on your own
machine. When you give agents fast, compiled tools written in Go, back them
with lightweight SQLite databases, and govern them with clear procedural
skills, you reclaim control over your environment.

You stop wondering why your API bill doubled. You stop dreading context
exhaustion. And you start treating coding agents for what they truly can be:
sharp, dependable partners in the craft of building software.

---

### The Stack at a Glance

For those who want to explore or run these tools locally:

- **Thermal:** Local-first agent FinOps and token tracking.  
  Repository: [github.com/jadmadi/thermal](https://github.com/jadmadi/thermal)  
  Install: `go install github.com/jadmadi/thermal/cmd/thermal@latest`  
  Project Overview: [/project/thermal/](/project/thermal/)

- **Skill Cabinet Go:** Census, audit, and deduplication for agent skills.  
  Repository: [github.com/jadmadi/skill-cabinet-go](https://github.com/jadmadi/skill-cabinet-go)  
  Upstream Standard: [subsy/skill-cabinet](https://github.com/subsy/skill-cabinet)

- **Agent Skills:** Curated procedural skills and verification harnesses.  
  Repository: [github.com/jadmadi/skills](https://github.com/jadmadi/skills)

- **Mahak Benchmark:** Open Arabic fluency benchmark and evaluation CLI.  
  Live Matrix: [mahak.waqf.dev/en/matrix/](https://mahak.waqf.dev/en/matrix/)

- **Profile & Catalog:** Overview of open-source tooling and standards.  
  GitHub: [github.com/jadmadi](https://github.com/jadmadi)  
  Projects Catalog: [/projects/](/projects/)
