In development · early-access list open

Release AI with evidence.

Evalora runs your AI feature against a repeatable suite of test cases every time you change a prompt or model, then shows you — case by case — what improved, what broke, and what it will cost.

  • qualityScore outputs against criteria you define, with flags for unsupported answers.
  • latencyTrack response time per case and across the suite.
  • costEstimate spend per run from token counts and the rates you configure.
case S-02 · shipping Demo data
inputDo you ship to Canada, and how long does it take?
source“Canadian delivery takes 5–8 business days.”
v1.4 ✓Yes, we ship to Canada. Delivery usually takes 5–8 business days.
v1.5 ✕Yes! We ship to Canada, and most orders arrive within 2–3 business days. ⚑ unsupported
Tone scoreimproved
Response timefaster
Groundingregressed

Fictional example — the new prompt reads better and breaks a fact.

Illustrative product preview

Two versions. One test suite. Every difference visible.

Switch suites, turn evaluation criteria on and off, filter to regressions, and open any case to compare outputs side by side. The pass/fail results, deltas and release rule update as you change the criteria.

Demonstration data. Every case, output, score, latency and cost here is fictional and stored in this page. No models are called and no live evaluation is running. Static demo
Evaluation criteria
Show
Baseline
Candidate
Case Input Baseline Time Est. cost Candidate Time Est. cost Change
Costs use illustrative per-token rates, not vendor pricing. The “demo release rule” is a simple example: any regression or new unsupported-answer flag means hold. In Evalora, a person makes the release decision.
The problem

A better answer here can mean a broken answer there.

AI features don't fail like ordinary code. A prompt tweak or model upgrade that fixes the case you were looking at can quietly change behaviour in cases you weren't. Most teams find out from users.

01 / inconsistent outputs

The same input doesn't always get the same answer

Model outputs vary between runs and between versions. A single spot check tells you how one response looked once — not how the system behaves.

run 1 → “5–8 business days”
run 2 → “about a week”
run 3 → “5–8 business days”
02 / manual testing

Testing lives in someone's spreadsheet

Teams paste a handful of prompts into a playground, eyeball the results and move on. It's slow, hard to repeat, and nobody can say later exactly what was checked.

prompts_to_try_FINAL_v3.xlsx
last updated: unknown
reviewer: unknown
03 / unexpected regressions

Changes ship without a before-and-after

Without a fixed test set and a side-by-side comparison, it's easy to release a change that improves tone and breaks facts, or cuts cost and drops key details.

tone ▲ · latency ▲ · cost ▼
grounding ▼ unnoticed
Planned capabilities

Built for teams who change prompts and models often.

Evalora is being designed around the regression-testing loop engineers already know — applied to AI outputs.

Test dataset management

Keep test cases, reference answers and source context in versioned suites. Tag by category, import from files, and grow the set from real failures.

Prompt and model comparisons

Run a baseline and a candidate over the same suite and see outputs side by side, with per-case quality, response-time and estimated-cost deltas.

Configurable evaluation criteria

Combine deterministic checks (format, length, required phrases) with rubric-based, model-assisted scoring. Choose which criteria gate a release.

Unsupported-answer flags

Highlight claims that aren't backed by the source context you supplied, so invented details get a second look before they reach users.

Human review queues

Route regressions, flags and low-confidence scores to reviewers. Their verdicts are recorded alongside automated scores and used to calibrate them.

Version history and release reports

Every run records the prompt, model, dataset version and criteria used. Export a release report showing what was tested, what changed and who signed off.

How it works

From test cases to a release decision.

  1. 1

    Add test cases

    Define inputs, reference answers and source context. Start with the cases you already check by hand.

    suite: support-returns@v3
  2. 2

    Run evaluations

    Run the current and candidate versions over the whole suite, recording outputs, timings and token usage.

    baseline ↔ candidate
  3. 3

    Compare results

    See pass/fail per criterion, regressions, improvements, response time and estimated cost — case by case.

    Δ quality · Δ time · Δ cost
  4. 4

    Review failures

    Reviewers inspect regressions and flagged answers, confirm or overturn automated scores, and add notes.

    queue → verdict
  5. 5

    Make a release decision

    Ship, hold or iterate — with a report that records the evidence behind the decision.

    release-report.pdf
Planned AWS architecture

Serverless orchestration, containerised workers.

Our current design for running evaluations at scale. It is a plan, not a description of a production system, and may change as we build.

Planned Evalora architecture on AWS The Evalora app starts a Step Functions workflow, which fans test cases out to ECS on Fargate workers. Workers read datasets from and write results to S3, and call Amazon Bedrock for supported models and model-assisted scoring. Results in S3 feed review queues and release reports. Evalora app & API Start runs, view results AWS Step Functions Workflow orchestration ECS on Fargate Evaluation workers run cases · apply criteria Amazon S3 Datasets & results versioned objects Amazon Bedrock Supported models + assisted scoring Review & reports Queues, release reports start run fan out cases invoke results cases read results escalate failures

ECS on Fargate compute

Containerised evaluation workers that execute test cases, apply deterministic checks and collect timing and token data. Scales with suite size without managing servers.

AWS Step Functions orchestration

Coordinates each run: splitting suites into batches, retrying failed calls, waiting on scoring, and marking the run complete.

Amazon S3 storage

Stores versioned test datasets, raw outputs and scored results, so any past run can be inspected and reproduced.

Amazon Bedrock models

Used for calls to supported foundation models and for model-assisted scoring against rubric criteria.

Region choices, data retention, encryption settings and support for models outside Bedrock are still being decided and will be documented before early access begins.

Methodology

Automated scores are indicators, not verdicts.

Automated evaluation makes it practical to check many cases on every change. It does not replace judgment. We treat every automated score as a signal that needs to earn trust against human review.

  • 01

    Prefer deterministic checks where possible

    Length limits, required fields, reply language and exact values can be checked with code. These are cheap, fast and repeatable.

  • 02

    Use model-assisted scoring for judgment calls

    Rubric-based scoring by a model can assess things like faithfulness or tone, but it can be wrong, biased toward certain styles, or inconsistent between runs.

  • 03

    Calibrate against human judgment

    Reviewers label a sample of cases. Comparing their verdicts with automated scores shows where a criterion can be relied on and where it needs a person.

  • 04

    Report what was — and wasn't — tested

    A test suite only covers the cases in it. Release reports list the suite, criteria and reviewer decisions so the evidence has clear limits.

Example calibration plan

Illustrative
CriterionScored byCalibration
Format & languageCode checkSpot-check rules
Answers correctlyReference + modelHuman sample each suite version
Supported by sourceModel-assistedAll flags human-reviewed
Follows instructionsModel-assistedHuman sample; re-check on rubric change
How much agreement is “enough” depends on the use case and the cost of an error. Teams set that threshold; Evalora is designed to show the comparison.
What testing can't do. Regression testing lowers the chance of shipping a known problem. It cannot guarantee an AI system is free of errors, and it won't catch failures that none of your test cases exercise.
About

An early-stage product, built in the open about where it stands.

Evalora is being built for product and engineering teams shipping AI assistants, AI features in SaaS products, and internal AI tools — teams who change prompts and models regularly and need a dependable way to know what changed.

We're in development. There are no customers, benchmarks or certifications to show yet, so this site doesn't claim any. We're looking for a small group of early-access teams to shape the product with us.

  • NowDesigning the dataset format, comparison views and evaluation runner.
  • NextPrivate early access with a small number of design-partner teams.
  • LaterHuman review workflows, release reports and wider model support.
Development statusIn development — pre-release
Product availabilityNot yet available; early-access list open
Live evaluations on this siteNone — the preview uses local demonstration data
InfrastructureAWS architecture planned, not yet in production
Company legal name[Company legal name]
Registered address[Registered address]
Company number[Registration number]
Founded[Year]
Founders / team[Names & roles]
Contact[hello@your-domain]
FAQ

Questions teams ask.

Can I use Evalora today?

Not yet. Evalora is in development. You can join the early-access list below and we'll contact you when design-partner places open.

Does the sample evaluation on this site run real models?

No. The preview uses fictional, hand-written demonstration data stored in the page. Filtering and toggling criteria recalculate results locally in your browser; nothing is sent anywhere and no model is called.

Will Evalora tell me whether my AI is correct?

It will show you how outputs perform against the test cases and criteria you define, and how that changes between versions. Automated scores are indicators to be calibrated against human review — not a guarantee of correctness.

Does testing eliminate AI errors?

No. Regression testing reduces the risk of releasing changes that break behaviour you've tested for. It can't catch failures your test cases don't cover, and AI outputs can still be wrong in production.

Which models will be supported?

The plan is to support models available through Amazon Bedrock first. Support for other providers and self-hosted models is under consideration; we'll confirm specifics before early access.

How is cost estimated?

From recorded token counts multiplied by per-token rates you configure. Estimates are for comparing versions and will not exactly match your provider's bill.

Where will my test data be stored?

The planned design stores datasets and results in Amazon S3. Regions, retention, encryption and access controls are still being finalised and will be documented before any customer data is accepted.

What does early access involve?

Early-access teams get to use pre-release versions, share a representative test suite, and give regular feedback. Terms and any pricing will be agreed individually. [Confirm early-access terms]

Early access

Help shape Evalora.

We're looking for teams who ship AI features and change prompts or models regularly. Tell us a little about what you're building and how you test it today.

  • Email: [early-access@your-domain]
  • Company: [Company legal name]
  • Location: [City, Country]

This form is not yet connected to a backend. [Connect form handler & privacy policy link]