From a single plain-language brief, the platform synthesises test scenarios together with the formal properties to be measured against them, drives optimization-guided exploration across millions of parameter combinations, and reasons about its own findings — converging on the regions where an automated-driving system is genuinely vulnerable, and returning not just the failing scenarios but a map of the design space they came from.
Generative AI is taking over scenario formulation itself. The engineer no longer specifies every candidate or every parameter — they ask for coverage of an operating domain, for corner cases, for stress on a controller. That is powerful, but it separates the engineer from the artifact by layers of automation. A tool that generates a scene and stops returns results without understanding — and when the system under test is itself an opaque learned policy with no code to audit, neither the scenario space nor the system is human-readable. Closing that gap is not a refinement. It is the only way an engineer can know what they have actually built.
Certification for higher autonomy implies scenario spaces of over a billion concrete tests. Exhaustive simulation is not merely expensive — it is structurally impossible. No engineer can mentally traverse it, and no fixed grid enumerates what matters.
Each software update — classical or neural — resets the verification clock and its compute bill. Software-defined vehicles evolve at software pace, so continuous re-verification is the baseline, not a one-off milestone.
Safety regulation for automated driving is still moving — and moving toward auditable, continuous, evidence-driven assessment. A defensible workflow must attribute findings to requirements as those requirements change.
Authoring, searching, and understanding are one continuous process. Each pass produces a better-aimed scenario than the last — and the return path is what generate-and-stop tools never build.
A plain-language ask becomes a valid, parameterised scenario in the open standard formats, together with the formal property to be measured against it.
Optimization-guided exploration maps where behaviour holds and where it breaks — returning the frontier, not a single verdict.
Each run is read as a designed experiment: statistical evidence, the requirement it implicates, and the refinement that should follow.
Authoring a scenario in the open standard formats is difficult. Authoring one that is meaningful — that genuinely exercises the property you care about — is harder still, and it is where most effort is lost before testing even begins.
Tempero turns a plain-language brief into a valid, parameterised scenario — and, alongside it, expresses the property to be measured on the ego trajectory in precise, formal terms. The two are distinct objects: a scenario stands on its own; a property is what validation and search evaluate against it. Getting them to align is the craft, and it is what makes the rest of the loop possible — without a precise statement of what must hold, there is nothing to search against and nothing to attribute a finding to.
This is where generation-only tools stop. It is where Tempero begins.
Classical optimization returns one optimum and discards everything else. Illumination — the quality-diversity paradigm — returns a structured archive: the most demanding case for every region of the space, and with it the limits themselves — beyond this speed, below this gap, past this curvature.
What comes back is not a scenario. It is a map: the critical regions, the gaps never explored, the representative extremes, and the parameters that genuinely matter. Because the system under test is often an opaque learned policy, this is how you use AI to make AI legible — the machine performs the exploration no engineer could, and the engineer keeps the judgement.
Illustrative: performance surface over a two-dimensional feature space
A meta-reasoning layer reads every completed run: it extracts the statistical evidence, ties each finding to the requirement it implicates, and proposes the next iteration of refinements — converging the search on genuine vulnerability rather than restating what was already known. Every proposal is provenance-tracked, and where statistical and regulatory guidance disagree, the conflict is surfaced rather than silently resolved.
That is one powerful way to use it. The same loop is equally an open-ended instrument: begin from a rough idea and refine a scenario until it surfaces the edge cases, near-misses, and limits nobody thought to look for.
The worst case found sits inside the requirement’s must-avoid envelope — a genuine violation, not a low-signal artefact.
The search ranks one parameter dominant, but the requirement names a different driver — inert here only because part of its range was never sampled.
A statistical refinement would prune a case the requirement names as limiting. Both are shown; the engineer decides.
The generated scenarios — valid, executable, and paired with a property specified precisely enough to measure.
An explicit representation of the space explored: its coverage, its structure, its boundaries, and the regions where behaviour turns critical.
Generate-and-stop delivers only the first. The more autonomy we hand to AI in design, the more the second becomes mandatory — and it is what Tempero returns by construction.
Not simulated benchmarks — validated against the actual test suites used for regulatory approval.
Addressing the “needle in a haystack” challenge of higher-autonomy certification. Tempero’s verification suite identified the only 2 critical safety cases out of over 40,000 valid scenarios — using a fraction of traditional computational resources.
Generated to explore the complete configuration space of the certification scenario
Simulation scenarios identified within regulatory constraints
Tempero conducted a detailed assessment of Automated Emergency Braking in two urban settings with varying traffic densities — demonstrating context-specific safety evaluation that global statistics cannot provide.
Setting td1 — AEB effective
Setting td2 — AEB limited
Unlike global statistics — results are tied to specific traffic and environmental conditions.
Assessment accounts for new mobility contexts including connected infrastructure
Standards-native, simulator- and compute-agnostic, and steerable in plain English through a standard agent interface.
Tempero (Latin: “to temper” or “refrain”) reflects our core mission: to mitigate the risks of high-risk AI systems. We identified a structural gap in the AI landscape long before the EU AI Act was finalized — a gap between academic legal frameworks and the rigorous technical proof required to ensure cyber-physical safety.
Our founding team brings deep R&D backgrounds in Automated Reasoning, Formal Testing, and Constraint Solving, with previous roles at Microsoft Research and leading autonomous systems institutes. We didn’t just build a tool for cars; we built a framework to translate high-level regulatory theory into executable safety evidence for AI decision-making in the physical world.
While our vision is broad, our execution is strategic. We prioritize Automotive SDVs today to leverage mature standards like OpenSCENARIO and meet the urgent global demand for scalable, certified safety.
Synthesis, search, and reasoning in one closed loop — returning the scenarios and the understanding behind them.
tempero.tech