Three non-negotiable engineering principles learned from running inference in high-stakes production. Each one is stated below with the shipped systems that hold it up — a principle nothing runs on is a preference, not a practice.
Unconstrained text generation has high entropy. Production systems require multi-pass schema validation, Pydantic guardrails, and deterministic state machines before any external tool executes side effects.
Throwing $20/token calls at trivial classification is lazy engineering. I prioritize quantization (AWQ/INT8), semantic caching, speculative decoding, and small specialized models on local runtimes.
If you cannot trace every token generation timestamp, p99 latency curve, prompt version, and tool-call parameter payload, your model isn't in production. It's an unmonitored experiment.
Multi-pass structured JSON extraction with Pydantic validation
Bounded tool recursion and recoverable checkpointed agent runs
Grounded writers that only claim what retrieval supports
INT8 quantization-aware training: 2× over FP16, 98.7% retained
Semantic caching and speculative decoding for p99 latency
Complexity-based routing that cut API spend by 65%
Per-token generation timestamps and prompt versioning in traces
p99 latency tracked per pipeline, not per model
Recovery from agent loops via Redis state checkpointing
I take on advisory and contract work where determinism, latency budgets, and cost governance actually matter.