Night circuit · agent evaluation Simulate · score · gate

Simulation · evaluation · CI gate

Prove the lap before the street.

Ovalen is a core AI/ML product: a simulation and evaluation platform for voice and chat agents. It runs thousands of staged callers, scores goal completion, and folds live misses into the next regression suite before a release lands. Ovalen does not ship the agent — it scores the one you already have.

  • Simulation staged callers
  • Evals goal + policy
  • CI gate on every release
Claude 3.7 SonnetGPT-4oGemini 2.5Llama 3 70BBedrockAnthropicOpenAI DirectGroqLangGraph Claude 3.7 SonnetGPT-4oGemini 2.5Llama 3 70BBedrockAnthropicOpenAI DirectGroqLangGraph

AI features

Where models sit on the oval

Ovalen is in the business of building a core AI/ML product for agent quality. Named jobs: staged-caller simulation, goal and policy scoring, production evals, and regression from live misses. Ovalen evaluates voice and chat agents. It is not itself an agent.

AI feature · Staged-caller simulation

Thousands of personas against voice and chat agents

For teams that need coverage before a release, not after a complaint

Graph-based personas with DTMF, noise, accents, policy traps, and tool failures. Import the agent over phone, WebSocket, or chat. Pin the version. Drive concurrent laps.

Persona pack → agent under test → POST /v1/runs

AI feature · Goal and policy scoring

Did the agent finish the job without breaking the rule

For quality leads who score accuracy, not just latency

Goal completion and policy sit on the same board as interruption and p95. A persona that passed last Tuesday can fail tonight and hold the deploy.

Transcript + tools → goal / policy scores → CI gate

AI feature · Production evals

Live calls sampled onto the same board as the suite

For operators watching the street, not only the paddock

Production evals on live calls: latency, interruption, goal, safety. Gray calls route to a person. Their labels become the next scorer.

Live sample → human or model scorer → next fixture

AI feature · Model-backed scoring

Claude, GPT-4o, Gemini 2.5, Llama 3 70B

For platform teams wiring LangGraph or a homebrew agent into CI

Planned model families for scoring and simulation: Claude 3.7 Sonnet, GPT-4o, Gemini 2.5, Llama 3 70B. Access: Bedrock, Anthropic, OpenAI Direct, Groq. Agent frameworks under test: custom/homebrew and LangGraph. Ovalen is not fine-tuning those models.

Suite turn → selected model → Bedrock / Anthropic / OpenAI / Groq

AI development tools / APIs. Speech recognition, chat, machine learning. MVP stage — no invented customers, funding, or uptime claims on this page. Ovalen scores agents; it does not replace them.

Method

How Ovalen ships

Pin a version. Drive personas. Hold the deploy when a metric misses.

01

Stage the oval.

Import your agent over phone, WebSocket, or chat. Pin the version.

02

Send the field.

Ovalen drives personas at the agent. You watch goal and policy scores.

03

Lock the gate.

Failing metrics hold the deploy. Passing suites print a readiness packet.

Stack

What you actually buy

A proving ground. Simulation, evals, review, and the loop back into CI — not a trained foundation model.

Simulate

Graph-based personas with DTMF, noise, and branching intents. Thousands of concurrent laps.

Observe

Production evals on live calls. Latency, interruption, goal, safety, all on one board.

Review

Route the gray calls to a person. Their labels become the next scorer.

Loop

A miss in June becomes a fixture in the suite. Releases inherit the street.

Start

A first lap in one afternoon.

Write ceo@ovalen.online. You get a test token, a public persona pack, and a 30-minute walkthrough on the agent you want to score.

See pricing