LLM Core/Research
Byte Pair Encoding implemented from scratch in pure Python, replicating tiktoken's core with byte fallbacks and a custom vocabulary tuned on technical corpora.
GPT-2/4 regex pre-tokenization, byte-level fallbacks, and a vocabulary retrained on technical text, trading pure-Python speed for a measurable drop in token count on domain corpora.
Tell me what you're building, the constraints you're working with, and where it breaks. I reply within a day.