AI features
Where models sit on the oval
Ovalen is in the business of building a core AI/ML product for agent quality. Named jobs: staged-caller simulation, goal and policy scoring, production evals, and regression from live misses. Ovalen evaluates voice and chat agents. It is not itself an agent.
AI feature · Staged-caller simulation
Thousands of personas against voice and chat agents
For teams that need coverage before a release, not after a complaint
Graph-based personas with DTMF, noise, accents, policy traps, and tool failures. Import the agent over phone, WebSocket, or chat. Pin the version. Drive concurrent laps.
Persona pack → agent under test → POST /v1/runs
AI feature · Goal and policy scoring
Did the agent finish the job without breaking the rule
For quality leads who score accuracy, not just latency
Goal completion and policy sit on the same board as interruption and p95. A persona that passed last Tuesday can fail tonight and hold the deploy.
Transcript + tools → goal / policy scores → CI gate
AI feature · Production evals
Live calls sampled onto the same board as the suite
For operators watching the street, not only the paddock
Production evals on live calls: latency, interruption, goal, safety. Gray calls route to a person. Their labels become the next scorer.
Live sample → human or model scorer → next fixture
AI feature · Model-backed scoring
Claude, GPT-4o, Gemini 2.5, Llama 3 70B
For platform teams wiring LangGraph or a homebrew agent into CI
Planned model families for scoring and simulation: Claude 3.7 Sonnet, GPT-4o, Gemini 2.5, Llama 3 70B. Access: Bedrock, Anthropic, OpenAI Direct, Groq. Agent frameworks under test: custom/homebrew and LangGraph. Ovalen is not fine-tuning those models.
Suite turn → selected model → Bedrock / Anthropic / OpenAI / Groq
AI development tools / APIs. Speech recognition, chat, machine learning. MVP stage — no invented customers, funding, or uptime claims on this page. Ovalen scores agents; it does not replace them.