Tweaks

Which End Do You Start From?

T1-first sounds like the responsible choice. It's usually the blind one. The generic fallback isn't just a capability. It's the instrument that tells you what to build next.

Everything described here was built, run, and broken on personal equipment, on personal time, against publicly available Pokémon TCG market data. The fallback traffic observations come off my own server. No employer systems, data, tooling, or internal processes are involved, and nothing here draws on them. The opinions are mine and are not offered on behalf of any employer.

the question that doesn’t fit

Which cards had their 10-day SMA cross above their 50-day in the last two weeks, above $100? On my system that’s a server-side screen. On a surface with no generic path, it’s answerable exactly if the designers shipped a screening endpoint with those parameters, and unanswerable otherwise. No fallback to degrade to. Run it the other way. If I’d built PokeScan starting from that screen, it would answer one kind of question: moving-average crossovers, in the windows and thresholds I chose. Cross-set comparisons, grading spreads, pack EV, anything else a user brought: not badly, not at all. That’s the T1-only ceiling: the designers’ imagination, moving only as fast as they ship, on whatever signal they have.

The promotion pipeline post left a question open. Fine, tiers and promotion, but which end do you start from? The tempting answer is the deterministic one. Fixed contracts, validated outputs, pre-built tools for the questions you anticipated, nothing improvised. Every argument in that post about reliability, cost, and the trust boundary points toward T1, so why not start there and skip the expensive middle?

Because it’s blind. Start from the generic end. Not because it’s cheaper; it isn’t. Because it’s the only end that can still teach you what to build. The rest of this is why, and the last clause above is where to look: what signal do they have?

T1-first blinds the pipeline

The entire promotion discipline runs on one input: observed fallback traffic. Every T3 invocation is a user spending their own tokens to tell you what your tools don’t cover.

Be precise about what it tells you. On PokeScan the fallback is a raw export, so what lands in my logs is the data slice the model asked for: which cards, which fields, which window. The procedure it ran on that slice happened in the conversation, and post 02 explains why I never see it. T3 traffic is a demand signal about data shapes. The analytical shape you infer. That’s still the highest-fidelity signal you’ll get: interviews and tickets approximate demand, and fallback traffic is a user’s own tokens spent on it.

A T1-only surface handles unmet demand two ways, depending on how far off the question is. Near miss: there’s an adjacent tool, the model calls it with parameters it won’t accept, validation fails, and the error lands in your logs. Noisy, but visible. Far miss: nothing adjacent, the model calls nothing, the user gets a polite no, and no artifact of the exchange survives anywhere you can see. You don’t just fail to answer. You fail to learn that the question existed. The far misses are where the roadmap hides.

[LOAD-BEARING] The generic fallback is not just a capability. It’s the instrument that measures what you should build next. Start T1-only with no intake and you’ve committed to guessing the question distribution.

the asymmetry that decides it

Compare the failure modes.

Start T3-first and your failures are loud and expensive: the bills, the latency complaints, the ninety-second queries. Annoying, visible, and information-rich, because every expensive query is a promotion candidate announcing itself. You pay tuition and receive a curriculum.

Start T1-first and your failures are silent and cheap: nothing errors, nothing costs much, and users quietly conclude the system can’t help them. You save the tuition and learn nothing.

[LOAD-BEARING] Expensive failure that teaches beats cheap failure that doesn’t.

That’s the whole decision. Start from the generic end, and make it as cheap as you can afford to be wrong in.

when T1-first is actually right

Two legitimate exceptions, stated precisely.

First, when the question distribution is genuinely known and stable. Some domains really do have twelve questions. If you’ve operated the manual version for years and the query log is a flat histogram over a fixed set, build the twelve tools and don’t apologize.

Second, when a generic path is unacceptable regardless of demand: write actions, irreversible operations, domains where an improvised answer is a liability event and not just a wrong number. The trust boundary argument cuts T1-ward here, and it should.

But even in the exception cases, steal the instrument. Log the refusals, and be honest about the mechanics, because a refusal only reaches your server if a tool got called. Near misses log themselves through validation. Far misses need somewhere to land: an intake tool whose entire contract is “nothing here fits,” whose only effect is a record, and whose description tells the model to call it when it has no fit. Every record carries what was attempted, what it meant, why it was refused, which tool was in play. The raw question is the one field you may not be allowed to keep. Fine. The other four are enough.

That’s a synthetic T3: demand measurement without the generic execution risk. It’s cheap, it works under the strictest governance posture, and it’s the one piece of the generic-first kit a T1-first team can adopt without changing anything else. It does depend on the model choosing to call it, which is the one part you don’t control. Make the tool description do that work. The expected failure signature of a deterministic surface without one is a roadmap assembled from support tickets, a quarter late.

the day-one kit

So what does starting from the generic end actually mean you ship first? Less than you think:

A raw export with field selection, so the fallback is as cheap as a fallback can be. An entity resolution stage, because nothing downstream works on ambiguous references. A health tool, so the model knows what the data can support. The intake tool, because some questions miss the export too. And logging on your side of every call.

That’s the whole day-one system. No screens, no technicals, no profiles. Those get built when the traffic tells you to, in the order the traffic tells you, which will not be the order you would have guessed. Mine wasn’t. The best change I shipped to PokeScan was a field selector on the raw export, not a tool, and it would never have been first on a whiteboard.

which end, then

The gradient isn’t a maturity ladder where T1 is the grown-up end and T3 is the prototype end. It’s a specialization spectrum with a learning instrument at one end and a set of hardened conclusions at the other. Start at the end that can still teach you something. Harden what it teaches. And if you must start deterministic, at minimum give the model somewhere to file what you refused. The questions you can’t answer are the roadmap. Silence is the only truly unaffordable failure mode.

phil-karpowich.com/writing/which-end-do-you-start-from