Agentic AI/Live · v2.0
Turns any GitHub repo into a live onboarding program. A nine-step autopilot parses the codebase into a dependency and entity graph, routes every model call free-tier-first, then turns what it finds into issues, tasks, PRs, and senior reviews.
Lead engineer and primary author: about 350 of the repo's ~370 commits, across the FastAPI backend, the autopilot pipeline, model routing and the React dashboard, with one collaborator.
A repository is ingested once into a dependency graph plus an entity graph of classes, functions and API routes, and that index is reused everywhere. The autopilot classifies issues by difficulty, assigns each one to whoever currently holds the fewest active tasks, opens a PR per issue, re-parses the PR head to graph-diff it against the base, and emits a structured senior review. Because a merged PR advances the linked task and closes the originating GitHub issue, the loop closes without anyone touching a board.
Keeping LLM cost near zero across a nine-step pipeline, and proving an autonomous PR actually fixed the issue rather than merely claiming it did.
Generated PRs are not trusted on the agent's word: step nine fetches each PR head, re-parses and re-graphs it, diffs the graph against the base for broken edges and new cycles, and retries unresolved issues a bounded number of times. Model choice was benchmarked on three real tasks from the repo across three OpenRouter models (9/9 solved; gpt-oss-20b about 5.6x cheaper than DeepSeek, gpt-4o-mini fastest at 6.8s average), and those fixes shipped as PRs with 51 passing tests across the touched suites. Backend and frontend CI run on every push.
The model benchmark is three tasks, which is enough to pick a default, not to rank models. Graph-diff validation catches structural breakage, not wrong behaviour that keeps the graph intact, so the senior review still assumes a human reads it. Free-tier routing trades latency and rate limits for cost.
Model routing is a cost-control problem before it is a capability problem: free tiers first with a per-query-type fallback chain, backed by exact and semantic Redis caches, kept a full autopilot run cheaper than a single paid call would have been. And a generated PR only counts once the validator re-parses it and diffs the graph, because the agent's own summary is precisely the thing you cannot trust.
Tell me what you're building, the constraints you're working with, and where it breaks. I reply within a day.