Skip to main content
Benchmarks

Benchmark Behaviour Runtime Intelligence on your agents

A practical evaluation rubric — not invented industry averages. Measure signal quality, situation readiness, and recovery discipline in your own environment.

Why we publish methodology, not vanity scores

Buyer trust requires measurable proof. Public “X% faster” claims without your topology are noise. These benchmarks define what to measure during a PUVINoise evaluation on app.puvilabs.com — and how scores should be interpreted by engineering, SRE, and leadership.

Dimensions

Five evaluation dimensions

Use all five in a single evaluation sprint. Record baseline (your current stack) vs PUVINoise path.

1. Signal path

Time from agent emit → visible in Command Centre. Pass: first trusted tenant-scoped trace within your agreed eval window.

2. Decision completeness

Share of runs with intent / tool-candidate / confidence fields. Pass: operators can reconstruct why a tool was chosen without log archaeology.

3. Situation readiness

Can a new on-call answer what is running, what changed, and what needs intervention from Command Centre alone?

4. Runtime Case MTTR path

Walk detect → triage → mitigate → learn once. Pass: case carries tenant, policy, and signal context end to end.

5. Tenant isolation checks

Verify tenant A signals never appear in tenant B views. Pass: security reviewer signs the isolation narrative.

Scorecard template

Rate each dimension 1–5 in your eval notes. Share with procurement alongside Pricing — not instead of hands-on proof.

How to run the benchmark

01

Baseline your current stack

Note how you detect agent intent failures today (APM, LLM tracer, tickets, chat). Capture one recent failure’s time-to-detect and time-to-recover qualitatively.

02

Instrument a sample agent

Install puvinoise-sdk, bootstrap, emit a traced run with tenant.id. Confirm ingest under the correct tenant.

03

Score the five dimensions

Use the rubric above with two operators (eng + SRE preferred). Record gaps honestly.

04

Compare commercially

Open Pricing for plan fit. Keep Engineering Consultation for scoped delivery — do not conflate product benchmarks with services estimates.

FAQ

Benchmark questions

Do you publish public latency percentiles?

Not as universal marketing claims. Latency depends on your network path, collector placement, and workload. We help you measure your path in evaluation.

How does this relate to LLM eval harnesses?

Offline/online model eval (scores, datasets) complements this rubric. Behaviour runtime benchmarks focus on production ops: situation, cases, and recovery. See the LLM Monitoring guide.

Where do case studies fit?

Case studies show qualitative operator journeys. Benchmarks tell you what to measure. Use both — then validate on app.puvilabs.com.

Linked from commercial paths

Product

Platform layers, Command Centre, Behaviour Intelligence, Runtime Cases.

Open Product

Run the benchmark on your fleet

Start a free evaluation — bring two operators and score the five dimensions in one session.