---
title: 'The Skill Bloat Dilemma: Why Your Coding Agent Can''t Handle 400 Skills'
description: >-
  As our agent skill libraries grew to hundreds of bundles, we hit a quiet
  cliff: silent alphabetical truncation, 50,000-token prompt bloat, and a 66%
  task failure rate. The architecture behind hybrid skill governance.
summary: >-
  Centralizing agent skills in global vaults creates a brutal trilemma:
  catastrophic prompt token bloat, silent alphabetical truncation past fifty
  tools, and the total blindspot of static per-project pinning. Here is how we
  solved it with grounded local catalogs, just-in-time MCP retrieval, and
  kernel-grade POSIX file locking.
date: 2026-09-29T00:00:00.000Z
heroImage:
  src: /_astro/hero.WuuY_PGO.webp
  width: 1360
  height: 765
  format: webp
tags:
  - agent-skills
  - token-optimization
  - ai-agents
  - context-window
  - mcp
  - systems-architecture
categories:
  - The Explainer
  - The Core Dump
published: true
featured: true
draft: false
author: Jad
href: /posts/the-skill-bloat-dilemma/
slug: the-skill-bloat-dilemma
---
A few weeks ago, I ran an audit on my local development environment.
Like many engineers leaning heavily into autonomous AI workflows, I had
gradually adopted the open convention of storing reusable procedural
instructions in `~/.agents/skills/`.

Whenever a teammate wrote a great workflow script, or someone published a
meticulous playbook for Cloudflare edge optimizations, or a community pack
for Go concurrency appeared, it went into the vault. It felt like the
responsible thing to do: build a permanent, compoundable library of
operational knowledge that any agent could draw upon.

Then I counted the directory.

There were 465 distinct skill bundles, comprising 3,188 individual
files and 21 megabytes of Markdown manuals, shell scripts, and JSON schemas.

I opened a fresh terminal window, initialized an agent session in a clean
repository, and looked at the token counter before typing a single prompt.

The system prompt was already hovering at **57,620 tokens**.

Before I had even asked the agent to inspect a file, it had consumed more
context than the entire source code of the project we were about to touch.
Every conversational turn, every clarification, and every tool call was
quietly dragging that massive fifty-seven-thousand-token anchor through the
API. Over a modest twenty-turn session, more than one million input tokens
were vanishing into thin air just to advertise capabilities we never once
touched.

We had built a library so comprehensive that it was choking the very
systems meant to read it.

---

## The Silent 50-Skill Cutoff Trap

My first instinct was to check how different harnesses handled this volume.
Surely modern tools like Google Antigravity, Gemini CLI, Claude Code, or
OpenCode had engineered an intelligent filtering mechanism for large vaults?

What I discovered under the hood was far more concerning than simple token
inflation.

To prevent prompt catastrophic bloat, several major harnesses implement a
hardcoded, unconfigurable ceiling: **they only advertise the first fifty
discovered skills to the model**.

At first glance, fifty sounds like a generous number. But how does an agent
engine select which fifty skills to keep when your vault contains 465?

It sorts them alphabetically by folder name.

Consider the mechanical reality of that choice:
1. In a 465-skill library, skills 51 through 465 (roughly 89 percent of
   your capabilities) are silently dropped before the prompt is ever
   assembled.
2. If you are working on a backend Go service, but your Go skills are named
   with the convention `golang-cli`, `golang-testing`, and
   `golang-database`, their alphabetical rank puts them somewhere around
   index 140.
3. As a result, your Go coding agent is given zero Go skills. Instead, it is
   handed fifty tools for accessibility audits, Docker scaffolding, and
   developer relations playbooks, simply because the letters A through D
   consumed the entire quota.

To verify this empirically, we constructed a controlled benchmark across
twelve realistic engineering tasks, ranging from standard flag parsing to
rare edge audits, evaluated over 108 trials.

Under default unmanaged discovery, the agent suffered an immediate
**66.7 percent failure rate**.

The failure was not caused by model reasoning deficits or bad code. The model
failed because the exact procedural guidance it needed had been amputated by
an alphabetical sorting rule.

---

## The Fragility of Static Manual Pinning

The obvious engineering workaround is to configure each project manually.
You create a local configuration file in your repository, say
`.agents/skills.json`, and explicitly pin the three or four skills that
matter for that specific codebase.

```json
{
  "skills": {
    "include": [
      "golang-cli",
      "golang-testing",
      "golang-database",
      "golang-error-handling"
    ]
  }
}
```

This immediately fixes the prompt bloat. The baseline catalog drops from
57,000 tokens down to roughly 500 tokens. Your agent launches instantly,
remains razor-focused on Go patterns, and your token invoices plunge.

For about forty-eight hours, you feel like you have solved the problem.

Then reality intervenes.

Software engineering rarely respects clean repository boundaries. You are
deep in the core Go CLI, and an urgent incident emerges: an edge worker
is exhausting its CPU time limits, or an SQLite migration script needs to
coordinate with a remote edge database, or you need to draft an API
announcement.

You prompt the agent to audit the issue. But because the project configuration
strictly pinned only Go tools, the agent is totally blind to your edge
worker skills.

In our 108-trial benchmark, static manual configuration scored a **complete
zero percent pass rate on off-stack and mixed tasks**.

Worse, maintaining static JSON manifests across dozens of repositories
introduces an exhausting maintenance tax. Every time a skill updates, or a
new teammate joins, someone has to manually curate, edit, and debug JSON
files across every project tree.

---

## The Skill Governance Trilemma

We found ourselves trapped in a three-way architectural tension:

```
                  Context Economy
               (Minimal Tokens / Turn)
                       /\
                      /  \
                     /    \
                    /      \
  Universal Capability ---- Zero Maintenance
 (Access to 400+ Tools)    (No Fragile Syncing)
```

1. **Default Global Discovery** gives you Universal Capability, but
   destroys Context Economy (57k tokens) and fails catastrophically via the
   silent 50-skill cutoff.
2. **Native Manual Pinning** gives you Context Economy (500 tokens), but
   destroys Universal Capability (zero access to off-stack tools) and creates
   high maintenance overhead.
3. **Crude Ad-Hoc Scripts** risk multi-agent file collisions, lost updates,
   and untracked configuration drift.

To solve this, we had to rethink procedural memory from first principles.
How does an operating system manage large programs with limited physical
RAM? It does not load every shared library into physical memory at boot. It
establishes a working set in L1 cache, and pages in remaining routines on
demand.

We needed a **Hybrid Skill Governor**.

---

## The Architectural Solution: Hybrid Governance

The design rests on three foundational pillars: **grounded local catalogs**,
**just-in-time retrieval**, and **kernel-grade atomic safety**.

```
+-------------------------------------------------------------+
| Host Global Vault: ~/.agents/skills/ (465 skills immutable) |
+-------------------------------------------------------------+
       |                                              |
       v [Read-Only Inspection]                       v [Dynamic JIT]
+-------------------------------+              +--------------------+
| Workspace: .agents/skills.json|              | Sila MCP Server    |
| (3-4 Grounded Domain Skills)  |              | 2 Tools: 230 tok   |
+-------------------------------+              +--------------------+
       |                                              |
       v [Active in Turn 1]                           v [On-Demand Only]
+-------------------------------------------------------------------+
|                     Model Context Window                          |
|  Base Prompt: 730 tokens (vs 57,620)  |  Fetch: Paid if invoked   |
+-------------------------------------------------------------------+
```

### 1. Grounded Local Catalogs
Instead of manually guessing which skills belong in a project, a read-only
assessment engine inspects the repository manifests (`go.mod`,
`package.json`, `wrangler.jsonc`, `Cargo.toml`) and active task goals.

It automatically pins the 3 or 4 highest-confidence domain skills directly
into `.agents/skills.json`. This reduces the baseline prompt overhead to
between 151 and 500 tokens on Turn 1, completely bypassing the native
50-skill cutoff.

### 2. The Just-In-Time MCP Bridge
To eliminate the off-stack blindspot, we expose two lightweight, read-only
Model Context Protocol (MCP) tools:
- `skill_search`: Allows the agent to query the entire 465-skill catalog by
  keyword or tag in under 2.5 milliseconds.
- `skill_load`: Retrieves full instructions and bundled reference resources
  dynamically into the conversational turn in under 300 microseconds.

Crucially, the tool declarations consume a fixed footprint of just **230
tokens**. The agent can reach any specialized tool in the global vault, but
we pay the payload cost *strictly on the turns when that tool is actually
called*.

### 3. POSIX Advisory Locking and Cryptographic Envelopes
In real agent environments, multiple background workers often run
concurrently. If two agents attempt to adjust configuration simultaneously,
uncoordinated file writes cause silent data loss.

We wrapped all harness writes in a two-tier protection model:
- **Kernel-level POSIX `flock`**: Enforces strict mutual exclusion at the
  OS level. If a process holding a lock is terminated with `SIGKILL`, the
  operating system kernel automatically releases the file descriptor in
  0.23 milliseconds, making deadlocks mathematically impossible.
- **Cryptographic SHA-256 Envelopes (`_sila`)**: Every managed file records
  a digest of its contents. If a human developer edits the file manually,
  the governor detects the hash mismatch, fails closed, and outputs a
  structured unified diff showing exactly which keys diverged.

---

## The Four-Part Cost Accounting Model

To measure the financial and attentional reality of this architecture, we
formalized the total context cost across any conversational turn $t$:

$$C_{\text{turn}}(t) = C_{\text{meta}}(t) + C_{\text{bridge}}(t) + C_{\text{fetch}}(t) + C_{\text{retained}}(t)$$

Where:
- $C_{\text{meta}}(t)$ is the catalog metadata for pinned workspace skills
  (151 to 500 tokens).
- $C_{\text{bridge}}(t)$ is the fixed 230-token declaration for the two MCP
  tools.
- $C_{\text{fetch}}(t)$ is the dynamic token payload delivered when a rare
  skill is loaded.
- $C_{\text{retained}}(t)$ represents instructions kept in conversation
  history across subsequent turns.

We mapped the amortization curves across both standard skills (2,700 tokens)
and heavy multi-file bundles (5,300 tokens):

| Conversational Turn | Unmanaged Global Vault | Native Truncated Cutoff | Hybrid Governor (Standard Skill) |
|---|:---:|:---:|:---:|
| **Turn 1 (Session Start)** | 57,620 tokens | 5,775 tokens | **730 tokens** |
| **Turn 2 (Rare Skill Invoked)** | 57,620 tokens | 5,775 tokens | **3,455 tokens** |
| **Turn 3 (Task Execution)** | 57,620 tokens | 5,775 tokens | **3,455 tokens** |
| **Turn 4 (Verification & Test)** | 57,620 tokens | 5,775 tokens | **3,455 tokens** |
| **Turn 5 (Commit & Handoff)** | 57,620 tokens | 5,775 tokens | **3,455 tokens** |
| **5-Turn Cumulative Total** | **288,100 tokens** | **28,875 tokens** | **14,550 tokens** |

For ordinary tasks, the savings are instantaneous: you save **4,757 tokens
on every single turn** compared to the 50-cutoff limit, and over 56,000
tokens compared to the unconstrained vault.

Even when fetching a massive 5,300-token bundle on Turn 2, the hybrid model
remains significantly cheaper than default discovery through Turn 8. And
unlike the truncated baseline, the hybrid model actually completes the task:

```
Overall Benchmark Accuracy (108 Controlled Trials):
  Default Global Discovery:    [████░░░░░░░░] 33.3% pass
  Native Manual Pinning:       [██████░░░░░░] 50.0% pass
  Sila Hybrid Governor:        [████████████] 100.0% pass
```

---

## When to Use What: A Pragmatic Mental Model

Not every codebase requires a full governance engine. As engineers, we must
avoid replacing one form of unnecessary complexity with another.

Here is the operational rule of thumb we adopted:

### 1. When Native Manual Pinning is Best
If you are working on a small, focused repository with a single programming
language, and your team relies on the same three or four static skills,
**just write `.agents/skills.json` by hand**.

You do not need an MCP server or dynamic retrieval. Avoiding the MCP bridge
saves you 230 tokens per turn, and your team avoids introducing any new
tooling dependencies. Keep it simple.

### 2. When Hybrid Governance is Mandatory
You need hybrid governance when:
- Your local vault exceeds fifty skills, and you are losing capabilities to
  alphabetical truncation.
- You maintain polyglot monorepos (e.g. Go backend, TypeScript edge worker,
  Python data scripts) where different directories require completely
  different active skills.
- Your workflows regularly encounter unforeseen, off-stack tasks that cannot
  be anticipated in advance.
- You run concurrent multi-agent swarms where uncoordinated file writes risk
  clobbering workspace configuration.

---

## Lessons from the Swarm

Building and verifying this system across three sequential goals taught us a
great deal about the emerging craft of multi-agent engineering.

Over the course of the project, more than forty autonomous subagents
collaborated across three sequential goals: an initial implementation team,
an empirical assessment panel, and a production hardening crew.

The most reassuring moment occurred during the hardening phase. An adversarial
challenger subagent was stress-testing flag combinations and identified an
obscure edge case: passing `--force --interactive` simultaneously caused a
subshell prompt bypass.

Rather than sweeping the issue aside, the gate **failed closed**. The
orchestrator halted forward progress, spawned a targeted remediation team,
restructured the flag precedence so that interactive prompts always take
priority, and verified the fix with unit regression tests before any release
candidate was authorized.

When you give agents clear boundaries, verifiable contracts, and independent
auditors who are incentivized to find flaws rather than rubber-stamp success,
they build remarkably resilient software.

---

## Conclusion

Agent skills are one of the most promising primitives in modern software
development. They allow us to capture hard-won operational wisdom in plain
text and hand it directly to our tools.

But like any powerful abstraction, they cannot be scaled naively. Dumping
hundreds of markdown files into a global folder and expecting a language model
to sift through them on every turn is the agent equivalent of disabling
virtual memory and running out of stack space.

By treating skills as a tiered memory hierarchy: keeping the working set
compact in L1 prompt cache, and paging in specialized tools via just-in-time
retrieval, we can preserve universal capability without paying for it on
every breath.

The code is pure Go. The binaries are standalone. The global library remains
untouched. And the terminal is quiet again.
