Reuse stable context
Prompt caching keeps repeated system instructions and shared context from being billed as fresh work on every call.
Independent AI engineer · available for the next hard problem
KostAI is a local-first LLM FinOps toolkit. It instruments each call, removes context that does not help, routes bounded work to cheaper compute, and keeps hard work on frontier models. I built it—and the systems below—to turn expensive AI ambiguity into measurable engineering proof.
Looking for AI platform, solutions engineering, and product engineering roles.
The original pipeline, rebuilt honestly
KostAI does not rely on one magic trick. It measures the call, removes avoidable input, reuses stable context, and routes only bounded work away from frontier models. The exact mix varies by workload.
Prompt caching keeps repeated system instructions and shared context from being billed as fresh work on every call.
Deterministic reducers remove duplicated history, boilerplate, filler, and other tokens that do not help answer the current task.
A local model extracts task-relevant facts while preserving exact values, identifiers, errors, and line numbers needed downstream.
Bounded, non-complex work can use a local or lower-cost model; hard or ambiguous work stays on the frontier path.
Protect the answer
Shadow mode tests a cheaper route beside the existing path while the production caller still receives the baseline result. That makes savings observable before a route is trusted.
The application behaves exactly as it did before while KostAI records cost, latency, and an output preview.
The proposed route runs in parallel. Its failure cannot replace or interrupt the baseline response.
Pass: candidate for controlled routing
Uncertain: keep or escalate to frontier
The 44.3% reduction is measured on a 41-task dogfood workload. KostAI has implemented shadow comparison and escalation safeguards; this page does not claim universal quality parity across every model, task, or production workload.
Pre-registered promotion protocol
Baseline and candidate outputs are stripped of route labels and scored from 0–10 for correctness, completeness, actionability, grounded confidence, and format fit. Preference, ties, and “both bad” outcomes are recorded before the route identity is revealed.
How the result was produced
The 44.3% figure is a workload-specific dogfood benchmark, not a universal promise. It compares naive Opus routing with KostAI’s routed execution across 41 tasks.
Record model, route, input/output tokens, cost, latency, and workflow without exporting prompt bodies.
Detect repeated context, oversized tool results, frontier overuse, and other recoverable token patterns.
Run the proposed route beside the baseline, compare outputs and cost, and keep the production answer unchanged.
Keep frontier models for complex work and hold routes that lack sufficient confidence or evaluation evidence.
Canonical definition
KostAI is John Bradley’s local-first LLM FinOps toolkit and portfolio anchor—not a model provider and not a universal savings guarantee.
Selected work
These projects connect the same operating pattern across cost, security, governance, and agent infrastructure: make the hard part observable, then automate it.
Local-first instrumentation, routing, and shadow-mode evaluation for finding expensive context and moving safe work off frontier models.
A self-improving operating system for prediction-market research and execution across Polymarket and Kalshi.
A zero-trust credential broker and verifiable evidence vault for AI agents, built around selective disclosure and provenance.
A threshold-cooperation primitive combining Pedersen commitments, FROST signatures, a hashchain, and an external verifier.
A dependency-free gateway between an AI agent’s intent and action, with explicit authority checks and synthetic safety gates.
Codebase-map compression and wire-format encoding for lower-cost coordination between coding agents through MCP.
What I bring
I’m most useful where engineering, product judgment, and a live customer problem collide.
Agent workflows, evaluation, model routing, MCP, and production guardrails.
React, TypeScript, Python, APIs, data pipelines, testing, and deployment.
Zero-trust identity, verifiable credentials, policy, and cryptographic protocols.
Turning ambiguous business pain into a working demo, proof, and implementation plan.
A no-regret first step
Start with a bounded shadow evaluation: keep every existing answer, measure avoidable spend locally, blind-review candidate outputs, and leave with a route-by-route evidence packet. No forced migration. No universal promise.
Discuss a shadow evaluation ↗The next build
If you’re hiring for AI platform, solutions, or product engineering work, I’d like to show you how I move from ambiguity to measurable proof.