Founding AI Engineer
Cekura (YC F24) builds testing and observability for voice and chat AI agents. I work on scenario generation, simulation pipelines and evaluation metrics.
- Architected an agentic scenario-generation system that turns an agent's prompt, tools and knowledge base into realistic test conversations covering tool-call flows, edge cases and KB-grounded Q&A, lifting scenario accuracy by 20–25%.
- Engineered a feedback-driven conversational refinement agent that improves scenarios through structured tool-calling and LLM reasoning, so customers iterate on tests in minutes instead of hours.
- Optimized the simulation pipeline for a 15% accuracy gain and a 50% latency reduction, and shipped an AI-powered end-of-call algorithm that cut premature or incorrect hang-ups to under 1% in production.
- Built agent-level evaluation metrics for intent alignment, response relevance, conversational consistency and interruption handling, then automated metric creation and doubled the speed of metric optimization.
- Developed customizable smart alerts that detect agent failures against client-defined metrics, enabling rapid diagnosis and recovery, plus red-teaming flows that probe what an agent should never say.
- Scaled production voice-AI systems that stress-tested 5M+ agent minutes, contributing to a 4× reduction in customer churn and 4× ARR growth in 8 months.
- Python
- LLMs
- OpenAI
- Gemini
- Vapi
- Deepgram
- Claude Code