Eval system audit
AI consulting · 2025Moved an LLM evaluation system to the Batch API (−50% cost) and removed nondeterminism: answer variance from 60% down to ~6% on an expert golden set.
AI Solutions Architect / Senior AI Engineer
Designing and delivering AI systems end to end: business-process analysis, solution architecture, PoC/MVP development, evaluation, infrastructure selection, deployment, and production iteration.
Current focus: AI agents, RAG, long-term memory, speech and meeting intelligence, developer tooling, and LLM evaluation. Products →
whytap · session recall · storm · selected consulting engagements

Led a team of 4 ML engineers and managed a 3-person delivery squad for the Construction AI project. Owned solution architecture, delivery, evaluation methodology, code review, and stakeholder communication across production AI projects.
Selected systems:
vLLM · Langfuse · Qwen-VL · Qdrant · E5

Owned pipeline health metrics (ingestion / freshness / context delivery), the evaluation framework, and agentic retrieval tools for an enterprise RAG platform over SharePoint (~12k sources) that unified departmental RAG pipelines. Case →
Azure ML Studio · Azure OpenAI · Azure Search · Ragas

Designed and delivered smart search over product composition with calorie filtering: 76% conversion to product, search time 54 → 35 s. Quantization shrank the model from 30 to 13 GB (generation 1.4 → 0.6 s). 60 RPS on 4× A100.
vLLM · Mistral · quantization · E5

Built the end-to-end agentic support pipeline at 2–6k tickets/day: auto-replies to ~20% of requests (99%+ confidence) cut load by 20%; operator suggestions brought a ticket from 1.5 min down to 37 s. Rolled out to 3 teams. Plus an AI meeting note-taker that files tasks into Jira.
LangChain · vLLM · Qwen · Milvus · faster-whisper
Side activity
Moved an LLM evaluation system to the Batch API (−50% cost) and removed nondeterminism: answer variance from 60% down to ~6% on an expert golden set.
Sized a GPU node for speech analytics (50k minutes / 30k calls a day): a node 30% cheaper, with headroom to scale to 100k minutes.
I teach senior developers to work effectively with agents. Trained 7 so far; two brought the approach into their own teams.