Blog
Notes from production.
Write-ups on LLM systems, agents and the unglamorous work that makes them reliable.
Retrieval isn't the problem. Ranking is.
How retrieval works in the Oracle of my D&D wiki: three search paths fused by rank, a reranker that only helps as long as it never sees the raw text, and leftover slots for what didn't fit.
Keeping the harness, swapping the model
What it takes to run Claude Code against NVIDIA NIM and local Ollama models when the subscription quota runs out and why the model that benchmarks best is not the one that works.