NEW

Replay any run from any step

Agents that finish
the job.

Norvane runs, traces and evaluates your agents in production, so every step is durable, replayable and explainable.

support-triage / runs /

run_8f2c1a

Live

S

support-triage

Runs

12.4k

Evals

38

Guardrails

6

Datasets

Alerts

API keys

Settings

Refund request #48213

Refund request #48213

Completed

10 steps

6.82 s

$0.0214

1 guardrail retry

Span

0 s

1.7 s

3.4 s

5.1 s

6.8 s

agent.run

llm.classify

412 ms

retrieve.policy

88 ms

tool.lookup_order

1.24 s

GET /orders/48213

llm.draft_reply

2.31 s

guard.pii_check

blocked

llm.redraft

1.87 s

guard.pii_check

tool.send_reply

Selected span

tool.lookup_order

Input

Output

Attributes

// arguments

{

"order_id": "48213",

"include": ["items", "payment"]

}

duration

1.24 s

attempts

2 (1 retry)

retry_reason

HTTP 503

checkpoint

ck_31a9

idempotent

true

Replay from here

support-triage / runs /

run_8f2c1a

Live

S

support-triage

Runs

12.4k

Evals

38

Guardrails

6

Datasets

Alerts

API keys

Settings

Refund request #48213

Completed

10 steps

6.82 s

$0.0214

1 guardrail retry

Span

0 s

1.7 s

3.4 s

5.1 s

6.8 s

agent.run

llm.classify

412 ms

retrieve.policy

88 ms

tool.lookup_order

1.24 s

GET /orders/48213

llm.draft_reply

2.31 s

guard.pii_check

blocked

llm.redraft

1.87 s

guard.pii_check

tool.send_reply

Selected span

tool.lookup_order

Input

Output

Attributes

// arguments

{

ʺorder_idʺ: ʺ48213ʺ,

ʺincludeʺ: [ʺitemsʺ, ʺpaymentʺ]

}

duration

1.24 s

attempts

2 (1 retry)

retry_reason

HTTP 503

checkpoint

ck_31a9

idempotent

true

Replay from here

Running in production at teams like these

oakline

oakline

Fennick

Quillon

Quillon

brightmoor

Harrow Labs

Harrow Labs

Solent

// mission

Agents are software. They deserve the same rigor as the rest of your stack.

Agents are software. They deserve the same rigor as the rest of your stack.

Agents are software. They deserve the same rigor as the rest of your stack.

Agents are software. They deserve the same rigor as the rest of your stack.

Logs you can replay. Retries you can trust. Tests that run before every deploy.

Logs you can replay. Retries you can trust. Tests that run before every deploy.

Logs you can replay. Retries you can trust. Tests that run before every deploy.

Logs you can replay. Retries you can trust. Tests that run before every deploy.

Norvane is the runtime that makes shipping an agent feel like shipping code.

Norvane is the runtime that makes shipping an agent feel like shipping code.

Norvane is the runtime that makes shipping an agent feel like shipping code.

Norvane is the runtime that makes shipping an agent feel like shipping code.

Everything between the prompt and production.

One runtime for the unglamorous parts of shipping agents: retries, state, visibility and proof that a change made things better.

Durable execution

Every step is checkpointed. When a model times out or a deploy restarts the worker, the run resumes where it stopped instead of starting over.

attempt 1

attempt 2

Resumed at step 9 from checkpoint ck_31a9. 8 steps not re-run.

Traces that read like a story

Every prompt, tool call and token, nested in the order it happened. Click any span to see exactly what the model saw.

agent.run

6.82 s

llm.classify

412 ms

tool.lookup

1.24 s

GET /orders

1.16 s

llm.draft

2.31 s

tool.send

470 ms

Evals on every deploy

Score each prompt or model change against your own dataset before it reaches users.

71

v11

78

v12

74

v13

89

v14

Guardrails in the loop

Block, rewrite or escalate a step before it acts. Policies live next to your code.

Redact card numbers

Refunds over $500 need approval

Block outbound links

Bring any model, any framework.

Norvane wraps the code you already have. No new DSL, no lock-in, OpenTelemetry out.

Python

TypeScript

Go

Local models

Any chat API

OpenTelemetry

Webhooks

Self-hosted

01 / Trace

See every step, not just the answer.

Every prompt, tool call and token, nested in the order it happened. Click any span to see exactly what the model saw.

Nested spans with cost and latency

Inputs and outputs per step

Share a failing run with one link

01 / Trace

See every step, not just the answer.

Every prompt, tool call and token, nested in the order it happened. Click any span to see exactly what the model saw.

Nested spans with cost and latency

Inputs and outputs per step

Share a failing run with one link

run_8f2c1a / trace

agent.run

2.1 s

llm.classify

412 ms

retrieve.policy

86 ms

tool.lookup_order

1.24 s

llm.draft_reply

2.31 s

guard.pii_check

34 ms

tool.send_reply

478 ms

run_8f2c1a / trace

agent.run

2.1 s

llm.classify

412 ms

retrieve.policy

86 ms

tool.lookup_order

1.24 s

llm.draft_reply

2.31 s

guard.pii_check

34 ms

tool.send_reply

478 ms

02 / Evals

Score every change before it ships.

Turn real failures into eval cases and gate each deploy on the result, so regressions fail the check and not your customers.

Datasets built from production runs

Code and model graders

A pass-rate gate in CI

02 / Evals

Score every change before it ships.

Turn real failures into eval cases and gate each deploy on the result, so regressions fail the check and not your customers.

Datasets built from production runs

Code and model graders

A pass-rate gate in CI

evals / refund-policy

94.2%

pass rate, 412 cases

71

v11

78

v12

74

v13

89

v14

94

v15

Deploy gate passed: 94.2% is above the 90% threshold

evals / refund-policy

94.2%

pass rate, 412 cases

71

v11

78

v12

74

v13

89

v14

94

v15

Deploy gate passed: 94.2% is above the 90% threshold

03 / Guardrails

Rules that run inside the loop.

Redact, block or escalate a step before it acts. Policies are code: versioned, tested and replayable like everything else.

PII redaction at the source

Spend and rate limits per run

Human approval for risky tools

03 / Guardrails

Rules that run inside the loop.

Redact, block or escalate a step before it acts. Policies are code: versioned, tested and replayable like everything else.

PII redaction at the source

Spend and rate limits per run

Human approval for risky tools

guardrails / production

Redact card numbers

412 hits

Refunds over $500 need approval

37 held

Block outbound links

9 blocked

Stay under $2.00 per run

0 hits

Cap tool calls at 40 per run

6 hits

Require approval for deletes

11 held

guardrails / production

Redact card numbers

412 hits

Refunds over $500 need approval

37 held

Block outbound links

9 blocked

Stay under $2.00 per run

0 hits

Cap tool calls at 40 per run

6 hits

Require approval for deletes

11 held

04 / Deploy

Ship on a Friday.

Durable execution checkpoints each step, so a deploy or a provider outage resumes the run instead of restarting it.

Resume from the last checkpoint

Canary by percentage of runs

Roll back in one click

04 / Deploy

Ship on a Friday.

Durable execution checkpoints each step, so a deploy or a provider outage resumes the run instead of restarting it.

Resume from the last checkpoint

Canary by percentage of runs

Roll back in one click

deploy / v15

Build image

18 s

Run 412 eval cases

41 s

Canary 5% of runs

6 min

Promote to 100%

12 s

Smoke test production

9 s

Notify #releases

1 s

Checkpointed runs resumed

0 lost

deploy / v15

Build image

18 s

Run 412 eval cases

41 s

Canary 5% of runs

6 min

Promote to 100%

12 s

Smoke test production

9 s

Notify #releases

1 s

Checkpointed runs resumed

0 lost

01 / 04 TRACE

02 / 04 EVALS

03 / 04 GUARDRAILS

04 / 04 DEPLOY

01 / Trace

See every step, not just the answer.

Every prompt, tool call and token, nested in the order it happened. Click any span to see exactly what the model saw.

02 / Evals

Score every change before it ships.

Turn real failures into eval cases and gate each deploy on the result, so regressions fail the check and not your customers.

03 / Guardrails

Rules that run inside the loop.

Redact, block or escalate a step before it acts. Policies are code: versioned, tested and replayable like everything else.

04 / Deploy

Ship on a Friday.

Durable execution checkpoints each step, so a deploy or a provider outage resumes the run instead of restarting it.

Three lines to instrument. Nothing to rewrite.

Wrap your existing agent, deploy as usual, and Norvane starts recording from the first request.

Instrument

Decorate your agent and tools. Norvane checkpoints each step automatically.

Observe

Set alerts on cost, latency or failure rate per agent, per customer.

Improve

Turn real failures into eval cases and gate every deploy on the score.

agent.py

python

1 import norvane

2

3 nv = norvane.init(project="support-triage")

4

5 @nv.tool(retries=3, idempotent=True)

6 def lookup_order(order_id: str) -> Order:

7 return orders.get(order_id)

8

9 @nv.agent(guardrails=["no-card-digits"])

10 def triage(message: str) -> Reply:

11 intent = classify(message)

12 order = lookup_order(intent.order_id)

13 return draft_reply(intent, order)

14

15 # every call is now traced and resumable

1 import norvane

2

3 nv = norvane.init(project="support-triage")

4

5 @nv.tool(retries=3, idempotent=True)

6 def lookup_order(order_id: str) -> Order:

7 return orders.get(order_id)

8

9 @nv.agent(guardrails=["no-card-digits"])

10 def triage(message: str) -> Reply:

11 intent = classify(message)

12 order = lookup_order(intent.order_id)

13 return draft_reply(intent, order)

14

15 # every call is now traced and resumable

Connected to support-triage

first trace in 41 s

Three lines to instrument. Nothing to rewrite.

Wrap your existing agent, deploy as usual, and Norvane starts recording from the first request.

Instrument

Decorate your agent and tools. Norvane checkpoints each step automatically.

Observe

Set alerts on cost, latency or failure rate per agent, per customer.

Improve

Turn real failures into eval cases and gate every deploy on the score.

agent.py

python

1 import norvane

2

3 nv = norvane.init(project=ʺsupport-triageʺ)

4

5 @nv.tool(retries=3, idempotent=True)

6 def lookup_order(order_id: str) -> Order:

7 return orders.get(order_id)

8

9 @nv.agent(guardrails=[ʺno-card-digitsʺ])

10 def triage(message: str) -> Reply:

11 intent = classify(message)

12 order = lookup_order(intent.order_id)

13 return draft_reply(intent, order)

14

15 # every call above is now traced and resumable

Connected to support-triage

first trace in 41 s

// the loop

Plan it. Run it. Prove it.

Three verbs cover the whole life of an agent. Each one leaves a trail you can read, replay and test.

01 / Plan

Plan it.

Plan it.

Plan it.

Plan it.

Give the agent a goal and tools. Norvane records the plan it drafts, so you can read what it intends to do before it does it.

Give the agent a goal and tools. Norvane records the plan it drafts, so you can read what it intends to do before it does it.

Typed tools

Dry runs

Plan diffs

agent.py

# the agent proposes, you review

plan = agent.plan(task)

# 1. lookup_order(order_id)

# 2. check refund policy

# 3. draft_reply(intent)

plan.approve()

02 / Execute

Run it.

Run it.

Run it.

Run it.

Every step is checkpointed. If a provider times out or you deploy mid-run, the run resumes from the last good step instead of starting over.

Every step is checkpointed. If a provider times out or you deploy mid-run, the run resumes from the last good step instead of starting over.

Durable steps

Retries

Idempotency

agent.py

run = agent.run(plan)

# step 9 timed out

# resumed from checkpoint ck_31a9

run.status # completed

run.cost # $0.024

03 / Audit

Prove it.

Prove it.

Prove it.

Prove it.

Replay any run, change one step and compare. Turn the failures into eval cases and block deploys that make the score worse.

Replay any run, change one step and compare. Turn the failures into eval cases and block deploys that make the score worse.

Replay

Eval gates

Audit log

agent.py

trace = run.trace()

fixed = trace.replay(step=4)

evals.assert_pass(fixed)

# 412 / 412 cases, gate passed

Replay

Rewind a bad run. Change one step. Play it forward.

What’s new in Replay

runs /

run_8f2c1a

/ replay

1 blocked step

Steps

6.82 s

1

Classify intent

llm

412 ms

2

Retrieve refund policy

retrieval

88 ms

3

lookup_order

tool

1.24 s

4

Draft reply

llm

2.31 s

5

PII guardrail

guard

34 ms

6

Send reply

tool

skipped

PII guardrail

type guard

model policy: no-card-digits

time 34 ms

tokens 0

cost $0.0000

Blocked

Input

"...to the card ending 4417."

Output

{
"blocked": true,
"rule": "no-card-digits",
"span": [58, 62]
}

Patch before replay

step 4, system prompt

- Confirm the refund and the card it was issued to.

+ Confirm the refund. Never repeat card or account digits.

Steps 1 to 4 reuse cached outputs.

// always on

Every step. Even at 3 a.m.

Every step. Even at 3 a.m.

Every step. Even at 3 a.m.

2.4B

steps traced a month

41 ms

p50 overhead

99.95%

runs resume after a failure

TRACE

EVALS

GUARDRAILS

LATENCY

Product tour / 1:32

DEPLOY

POLICY

ALERTS

REPLAY

// the demo

See the whole run.

THE THING IS

Agents that ship,

Agents that ship,

not agents that chat.

not agents that chat.

A demo can be faked.

A demo can be faked.

A replay can’t.

A replay can’t.

We don’t build for the demo.

We don’t build for the demo.

We build for the 3 a.m. page.

We build for the 3 a.m. page.

THE THING IS

Agents that ship, not agents that chat.

Agents that ship, not agents that chat.

Agents that ship, not agents that chat.

A demo can be faked. A replay can’t.

A demo can be faked. A replay can’t.

A demo can be faked. A replay can’t.

We don’t build for the demo. We build for the 3 a.m. page.

We don’t build for the demo. We build for the 3 a.m. page.

We don’t build for the demo. We build for the 3 a.m. page.

Built for the traffic you are about to have.

2.4B

agent steps executed on Norvane last month

2.4B

agent steps executed on Norvane last month

41

ms

median overhead added per traced step

99.95

%

uptime commitment on Team and Enterprise

6.3

x

faster root cause, as reported by customers

We stopped guessing why agents failed. Norvane shows the exact step and input, and we replay the fix in minutes.

We stopped guessing why agents failed. Norvane shows the exact step and input, and we replay the fix in minutes.

ML

Mara Lindqvist

Head of AI Platform, Oakline

oakline

Durable runs paid for it in week one. A provider outage used to mean re-running 40-minute jobs from zero.

TV

Teodor Vance

Staff Engineer, Fennick

Durable runs paid for it in week one. A provider outage used to mean re-running 40-minute jobs from zero.

TV

Teodor Vance

Staff Engineer, Fennick

Evals on every deploy turned prompt changes from a gamble into a normal code review.

AB

Aiko Brennan

CTO, Quillon

Evals on every deploy turned prompt changes from a gamble into a normal code review.

AB

Aiko Brennan

CTO, Quillon

  • Python

  • TypeScript

  • Go

  • OpenTelemetry

  • Webhooks

  • MCP servers

  • Local models

  • Any chat API

  • Self-hosted

  • Postgres

  • Kafka

  • Slack

  • Python

  • TypeScript

  • Go

  • OpenTelemetry

  • Webhooks

  • MCP servers

  • Local models

  • Any chat API

  • Self-hosted

  • Postgres

  • Kafka

  • Slack

  • Agents that ship

  • Agents that ship

  • Agents that ship

  • Agents that ship

  • Agents that ship

  • Agents that ship

  • Agents that ship

  • Agents that ship

  • Agents that ship

  • Agents that ship

  • Agents that ship

  • Agents that ship

Start free. Pay per step in production.

No seat taxes. Your bill grows only when your agents do real work.

Compare all plans

Hobby

$0

forever

For prototypes and side projects.

25,000 steps per month

7-day trace retention

1 project, 2 members

Team

Most teams

$49

per month

For agents with real users behind them.

250,000 steps, then $0.40 per 1k

Durable execution and Replay

Evals, guardrails and alerts

Enterprise. SSO, VPC or self-hosted deployment, and a 99.95% SLA.

Questions, answered.

Something missing? Talk to sales and an engineer will reply within a day.

What counts as a step?

A step is one model call, tool call or guardrail check inside a run. Retries of the same step are free, and cached steps during Replay are never billed twice.

Do I need to rewrite my agents?

Which models and frameworks are supported?

Where is my data stored?

Can we self-host Norvane?

What happens if we go over our plan?

Put your first agent in production this week.

Free up to 25,000 steps a month. No credit card, no sales call.

$

pip install norvane

Create a free website with Framer, the website builder loved by startups, designers and agencies.