GSP / BENCHMARKS & METHODOLOGY

A result.
With its full context.

Inputs, quality checks, source identity and execution environment travel with the measurements.

Measured execution.
Not imagined performance.

5 repeated bounded engine executions per synthetic scenario. Wall time includes Python subprocess startup. Independent checks validate the stated outcome. No competitor workloads, hardware PPA or customer ROI are compared.

ScenarioVerdictRepeatsMedian wall timeStable artifact identities
Feasible baselineSynthetic input / known-answer checkcompleted565.398 msYes
Impossible failure-domain diversitySynthetic input / known-answer checkblocked526.28 msYes
Working set exceeds selected memorySynthetic input / known-answer checkblocked554.351 msYes
Insufficient power budgetSynthetic input / known-answer checkblocked524.923 msYes

Linux x86_64 / Python 3.13.5. Includes process startup. These are local software timings, not silicon PPA, customer ROI or a competitor comparison.

Compare the work.
Not just the stopwatch.

A blocked case can be the correct answer. The benchmark retains impossible diversity, insufficient memory and constrained power alongside feasible cases.

01

Exact scenario and engine hashes

02

Independent outcome checks

03

Repeated local execution

04

Raw artifacts and clear boundaries