Tweaks

The Promotion Pipeline

Generic AI capability and specialized tooling are in active tension. You don't resolve it by picking a point on the curve. You resolve it in time.

Everything described here was built, run, and broken on personal equipment, on personal time, against publicly available Pokémon TCG market data. The token counts and latency observations come off my own server. No employer systems, data, tooling, or internal processes are involved, and nothing here draws on them. The opinions are mine and are not offered on behalf of any employer.

Every team building AI-on-data systems eventually hits the same wall, usually without naming it.

You start with something generic: an export tool, a raw query capability, a “just give the model the data” fallback. It’s magic. Users ask questions you never anticipated and get answers. Then the bills arrive, the latency complaints arrive, and somebody asks why the same question takes ninety seconds and costs real money every single time it’s asked.

So you build specialized tools. They’re fast, cheap, reliable. And immediately less flexible. Now users hit walls the generic path never had, and you’re back where you started, except with more code to maintain.

This is not a tuning problem. It’s a structural tension, and I think there’s one pattern that survives contact with it. I call it the promotion pipeline.

the tension, stated properly

Generic capabilities are valuable because they answer questions you didn’t anticipate. Specific tools are valuable because they’re cheap, fast, and reliable. These are not different operating points on one curve you can optimize. They’re in active tension.

Every hour you spend hardening a specific tool is a bet that the question will recur in that exact shape. Every generic capability you preserve is a bet that novelty will keep showing up. Both bets are sometimes right. Neither is right permanently, because the question distribution itself moves as users learn what the system can do.

[LOAD-BEARING] You cannot resolve the generic-versus-specific tradeoff spatially, by picking a point. You can only resolve it temporally, by moving individual workloads along the curve as they prove themselves.

grounding: three tiers on a real system

PokeScan is my testbed for this: an MCP server over Pokémon TCG market data, tens of thousands of instruments, full price history, grading data. (The intro post covers why that market; short version, real structure, toy stakes.)

Its capabilities sit on a specialization gradient. The gradient is not about where the code runs. It’s about how much of the procedure the model is allowed to invent.

T3, improvised capability. pokescan_export will hand the model any slice of raw data and let it figure the rest out. Generic primitives, model-written code, whatever the question seems to need. Maximum flexibility. The model decides how the work gets done.

T2, reusable skills. A predefined analytical procedure (moving averages, screens, comparisons) that the model selects and configures by constructing parameters. On PokeScan those procedures live server-side, but that is one implementation pattern, not the definition. The same skill running in the agent’s sandbox is still T2. The model chooses the procedure and supplies the parameters. It no longer invents the procedure from scratch.

T1, hardened tools. One question shape, narrowed to a stable contract: constrained inputs, defined outputs, real tests behind it. Entity resolution, a specific screen, a specific profile. Nearly deterministic in and out.

Same question, different tiers, wildly different economics. A representative multi-card technical analysis on the T3 path burns roughly 95K tokens; the T1 path answers it in roughly 43K. That’s roughly 2.2x the input volume. Same runs, the T3 path was also materially slower: more model-side work, more tool round trips. Treat the latency as directional, not a benchmark. Reliability drops too, since every token the model touches is a chance to touch it wrong.

why the naive answers fail

“Just build the specific tools upfront.” You don’t know the question distribution yet. You’ll build the wrong tools, and you’ll pay the design cost for questions nobody asks.

“Just eat the generic cost.” The fallback is where volume lives (more on this below), so you’re eating the cost precisely where it compounds fastest.

“Cache it.” Caching helps identical queries. Analytical questions are near-identical in shape and different in parameters. The cache hit rate on “same shape, new card, new window” is roughly zero.

the mechanic

The promotion pipeline is an operating discipline, not a feature. Improvised workload, codified reusable skill, hardened product capability, in that order:

Watch T3. The fallback is your discovery instrument. Every T3 invocation is a user telling you, with their own tokens, what question shape your specific tools don’t cover.

Codify into T2. When a shape recurs, write the computation down once and expose it as a parameterized skill. Server-side is how I do it; a skill file the agent loads into its own sandbox counts too. The model stops improvising the math and starts filling in the blanks. This is the cheapest promotion: one function, no new tool surface.

Harden into T1. When a T2 skill’s parameter distribution stabilizes, collapse it into a purpose-built tool. Fixed contract, tight schema, boring on purpose.

Observed usage is the trigger. Design-time judgment sets the initial safety envelope; after that, the traffic decides. The decision-maker is whoever operates the server, looking at real T3 traffic. That’s the minimum viable promotion loop, the part that has to exist before anything more elaborate matters.

the part nobody expects: promotion is your observability strategy

Here’s the constraint that turns this from optimization into necessity, and it’s a property of the protocol rather than a choice anyone made.

An MCP server never sees the conversation. What arrives at my server is a tool name and a parameter dict. Not the user’s question, not the model’s reasoning, not the context that shaped the call. I own the server, the pipeline, and the databases, and I still can’t see any of that, because it never crosses the boundary. There is no logging setting that would give it to me.

So the layer where the most consequential decisions happen (what the question actually meant, which tool to reach for, what parameters to construct) is structurally invisible to the layer I control. That’s not a gap in my instrumentation. It’s where the protocol draws the line, and every MCP server operator inherits it.

You can’t fix that with better logging. What you can do is shrink it. Every workload promoted out of T3 takes a procedure the model was improvising inside the conversation and pins it to something written down. A T1 tool has a contract you can regression-test. A T2 skill has a fixed procedure and a parameter set you can audit. Where that skill executes decides how much of the run lands in my logs, and running it on my own server is worth having. But the part that shrinks the blind spot is that the procedure stopped changing per query. The model still decides which capability to call. It stops deciding how the math gets done.

[LOAD-BEARING] The promotion pipeline is the observability strategy. You don’t get to watch the model think, so you systematically reduce how much thinking you’re asking it to do.

That reframes promotion as risk reduction, not just cost reduction. The cost savings are the incentive. The shrinking blind spot is the point.

the counterintuitive ROI

Instinct says optimize your best path: make T1 faster, add more T1 tools. Wrong instinct.

The fallback is where volume lives. By construction, T3 catches everything your specific tools don’t, and early in a system’s life that’s most of the traffic. So the highest-leverage engineering change is usually to the generic path itself. On PokeScan, the single best change I ever shipped wasn’t a new tool. It was a fields parameter on the raw export, letting the model request only the columns it needed. One parameter, and it beat every purpose-built tool I’d added, because it cut cost on every uncovered question at once.

the serious challenger, briefly

One more thing before closing, because the sharpest current objection deserves acknowledgment even though it gets a full post of its own: code execution approaches (the agent writing scripts in a sandbox instead of making tool calls) have made the generic fallback radically cheaper. That changes the economics above without touching the argument, because cost was the symptom and model-improvised-per-query is the property. The tiers come through it intact, because they were never about where code ran. An agent writing one-off code for this request is T3, cheaper T3, still T3. An agent calling a reusable skill somebody wrote and tested is T2, sandbox or server. A narrow, hardened interface is T1. A later post takes the challenger apart properly.

what generalizes

Nothing above is about Pokémon. The pattern applies to any system that serves open-ended natural language queries over a bounded set of data sources, needs to balance flexibility against cost and latency, and can observe its own usage. Which is to say: every production MCP server, and most agent systems, whether their operators have noticed yet or not.

Watch the fallback. Codify what recurs. Harden what stabilizes. Optimize the generic path first, because that’s where the volume is. And do all of it knowing the promotion itself is how you claw back observability the protocol was never going to hand you.

The next post answers the question this one raises for anyone starting fresh: fine, tiers and promotion, but which end do you build first?

phil-karpowich.com/writing/the-promotion-pipeline