Things built to understand how they work, not to ship next quarter, a vector database from scratch, GPT-2 from first principles, a tokenizer in pure Python.
A full vector database from scratch: 9 ANN algorithms (HNSW, IVF, PQ, Int8, LSH, KD-Tree, VP-Tree, BM25, Hybrid RRF), a cost-based query planner, WAL with fsync durability, background compaction, distributed scatter-gather, and row-level RBAC. Multimodal ingestion, cross-encoder RAG, OpenAI-compatible endpoints, and GPU-accelerated indexing via CuPy and Numba.
Reimplementing GPT-2 from first principles following the nanoGPT curriculum, character-level tokenization, causal self-attention, multi-head attention, the full transformer block, then byte-pair encoding and training on Tiny Shakespeare.
Architecture for recursive state persistence in non-deterministic agent clusters. Episodic memory buffers with decay-aware consolidation, semantic compression of long-running agent state, and checkpoint-based recovery.
Pure Python byte-pair encoding replicating tiktoken's core, with GPT-2/4 regex pre-tokenization, byte fallbacks, and a custom vocabulary trained on technical corpora.
Standardized handshake and task-routing protocol for multi-agent systems. Capability advertisement, task decomposition contracts, result aggregation, and failure escalation across heterogeneous runtimes.
Memory compression that prioritises relevant context for long agent sessions. Semantic chunking with recency-relevance scoring, sliding-window eviction, and compressed summaries to stay under budget without losing signal.
Reimplementing a vector index or a transformer block is the fastest way to find the assumptions a library hides from you. The knowledge transfers directly into the systems I ship for clients.