William White
Forward Deployed Engineer · Tampa, FL
I deploy agent systems into other people's live operations, usually on hardware they own, and I measure whether they actually work before I claim they do.
- 0%
- fabrication rate
- Across 51 held-out decisions, graded against an expert-annotated set the agents could not query. Inter-rater kappa 0.95.
- $0.00
- decision-path API spend
- The whole pilot ran on Hermes-4-70B-FP8 under vLLM on one H100, network isolated. Total compute was about $8.50.
- 600/600
- composes returned ok
- Zero parse failures across two model families (321 on an 8B local tier, 279 on a 70B tier).
- 2 of 4
- hypotheses refuted
- Pre-registered before the run, published anyway when the data went against me. Reproduces bit for bit.
What I do
Systems other people run their business on, in their environment, every day.
- ·OTTER: 51 tables, 143 API routes, 35 permissions, in daily use at a 30-employee contractor managing 40+ active jobs
- ·Shafer Law Payments: cards and ACH under PCI SAQ-A, integer-cent accounting, append-only audit log
- ·Little Bear Foundry: nine role-specialized agents beside a customer's live operation, every action human-approved
Pre-registered hypotheses, frozen holdouts, and the findings that went against me published with the rest.
- ·Extended thinking cut LangGraph recall from 1.000 to 0.600 and raised Vercel precision from 0.714 to 0.938. Same model, opposite behavior by orchestration.
- ·Graph-only retrieval refused 100% of adversarial distractors where hybrid refused 10%
- ·Three published studies, two of them with refuted hypotheses in the results table
Open-weight models on hardware the customer owns, and the MCP servers and clusters that hold it together.
- ·backbone-on-k8s: NetworkPolicy enforcement verified under Cilium, $0.00767 per task across 1,111 sessions
- ·sparql-mcp: every SPARQL write verb rejected before it reaches the store, host-guarded to localhost
- ·causal-evidence-mcp: returns "insufficient" below a minimum cohort size instead of a number nobody should trust
I work at Bay West Labs, which drops engineers and agent systems into mid-market operators, on-prem when the data cannot leave the building. Across client engagements the median from kickoff to first working build is about 14 days, all code and documentation transferred, and no system abandoned after handoff.
Everything on this site links to a repo, a PDF, or a running demo. If a claim here cannot be checked, tell me and I will cut it.
Sheldonwhite888@gmail.com · Resume (PDF) · GitHub