Orbtrace
On-prem only— your servers, your data, no SaaS

Self-hosted observability.We tell you the cause.

Orbtrace is the self-hosted observability platform built for SRE teams. It unifies logs, traces, and metrics with Causal RCA and Time-Travel Replay — powered by the LLM you pick: Claude, OpenAI, Azure OpenAI, Google Gemini, or a local Ollama. Runs entirely inside your network. OpenTelemetry-native. Air-gapped friendly. Docker Compose, Helm, or Ansible — your platform team's call.

Custom quote— unlimited hosts, no metering
Compose · Helm · AnsibleOn-prem onlyAir-gapped friendlyStarts on an 8 GB VPS (Eval)
~/orbtrace
zsh
$ orbtrace pilot init --license pilot-aug26.lic
→ docker compose up -d
 ✓ doris-fe started in 6.2s
 ✓ otelcol up · ports 4317, 4318
 ✓ orbtrace-server ready · http://localhost:8080
▸ Point your OTLP exporters at otel://localhost:4317
$ orbtrace query 'level:Error AND duration:>500'
orbtrace.your-company.internal/incidents/i-9f3
v1.0.0
level:Error AND service:checkout-api AND duration:>500 in "last 24 hours"128 hits · 92ms
Errors / min
312+184%
p99 latency
1.24s+412 ms
Ingest rate
128k/sstable
Errors over time
312 errors · checkout-api · last 24h
anomaly · 12:34
Trace · 7c91f4… · 1.24s · 7 spans
  • checkout-api · POST /orders1240ms
  • payments · charge970ms
  • stripe.charge (external)740ms
  • auth.verify (jwt)70ms
  • ledger · acct.lock120ms
  • postgres.tx (deadlock)100ms
  • kafka.publish orders.created50ms
Service topology
Causal RCAconfidence 0.91grounded · 312 cited spans

Root cause: deploy svc-payments@a91c2 introduced a long-running ledger.acct_lock tx. p99 across the checkout path increased by +412 ms. Rolling back resolves the incident.

Built on the open standards your platform team already runs

  • OpenTelemetry-native
  • Apache Doris
  • Docker Compose
  • Helm chart
  • Ansible playbook
  • GDPR · SOC 2 II
  • Runs on 8 GB RAM
  • OIDC · SAML · SSO
  • Air-gapped install
  • Anthropic Claude
  • OpenAI
  • Azure OpenAI
  • Ollama local
  • OpenTelemetry-native
  • Apache Doris
  • Docker Compose
  • Helm chart
  • Ansible playbook
  • GDPR · SOC 2 II
  • Runs on 8 GB RAM
  • OIDC · SAML · SSO
  • Air-gapped install
  • Anthropic Claude
  • OpenAI
  • Azure OpenAI
  • Ollama local
100k+
events / sec on a 3-node cluster
< 60s
from docker compose to first ingest
0%
error drop rate during incident peaks
8 GB
minimum RAM (default stack)
One backend

Three pillars, one query, zero pivot-tab fatigue.

Logs, traces, and metrics live in the same OLAP store (Apache Doris). One query language. One UI. No more 'wait, which tab had this trace?'

  • ERRledgerdeadlock postgres tx=acct.lock
  • WRNpaymentsstripe.charge p99 1.2s
  • INFcheckoutPOST /orders 202 84ms
  • INFingestdoris flush rows=128k
Logs

Search 100M lines in milliseconds

Native inverted indexes (Apache Doris) deliver 3–10× faster full-text search than ClickHouse in our benchmarks. Drain-style template clustering surfaces anomalies (NEW · SPIKE · DROP) so you stop scrolling raw lines. Lucene-style key:value syntax — the same shape Kibana, Datadog, and Loki use.

Learn more
Traces

No orphan spans, no manual stitching

Async/queue context reconstruction across Kafka, RabbitMQ, SQS, Redis Streams, NATS. Stitched parents shown with confidence scores you can trust.

Learn more
Metrics

No custom-metric tax — ever

Cardinality is a feature, not a billing line. Add as many tags as you need. Per-cluster license, predictable spend, no surprises at renewal.

Learn more
Use cases for SRE teams

Built for the on-call rotation, not the slide deck.

Eight workflows where Orbtrace meaningfully changes how SRE teams operate. Pick the ones that hurt today.

Incident response

Page-to-cause in minutes, not hours. Causal RCA names the suspect deploy, the upstream lock, the canary that tipped over — with span citations the on-call can verify.

  • MTTR
  • p99 latency
  • error budget
MTTR vs prior stack−74%

SLO tracking

Every service ships with multi-window, multi-burn-rate alerting (Google SRE handbook style). Burn-rate panels live next to the trace waterfall — no tool-switching.

  • SLI
  • SLO
  • burn rate
fastest burn detected14.2×

Post-mortem & review

Time-Travel Replay generates a counterfactual narrative: what would have happened without the deploy. Drop the result straight into your blameless post-mortem doc.

  • counterfactual
  • evidence
  • audit-ready
shorter post-mortems10×

Capacity planning

Aggregate p95/p99 latency and ingest-rate budgets per service. Cost-aware sampling exposes a real-dollar figure — not events/sec — so finance can sign off.

  • capacity
  • $ budget
  • tail sampling
custom-metric tax$0

Migration off Datadog & New Relic

Point your existing OpenTelemetry SDKs at Orbtrace's OTel Collector. Keep your dashboards and alerts. We import Datadog metrics names where it makes sense.

  • OTel-native
  • no agent rewrite
  • drop-in
to first ingested span60s

Audit & compliance

Synchronous deletes, immutable audit log, RBAC, and per-tenant data partitioning. GDPR Art. 17 erasure and SOC 2 evidence collection are built in, not bolted on.

  • GDPR
  • SOC 2 II
  • HIPAA-ready
data stays on-prem100%

Pre-deploy verification

Hook Orbtrace into your CI: every PR gets a synthetic replay against last week's traffic. Reject deploys that would have spiked p99 — before they reach prod.

  • CI
  • pre-deploy
  • synthetic replay
rollback-only deploys−83%

Async / event-driven

Stitch context across Kafka, RabbitMQ, SQS, Redis Streams, NATS. Confidence-scored stitched parents tell you which spans likely belong to the same business transaction.

  • Kafka
  • RabbitMQ
  • SQS
  • NATS
orphan spansminimized
On-premise only — by design

Built for teams who cannot ship their telemetry to a third party.

Datadog, HyperDX-Cloud, New Relic — beautiful products that solved visibility for the public-cloud era. They're also a non-starter the moment your auditor reads the data-flow diagram.

Your data never leaves your network

Telemetry — every span, every log line, every metric — is ingested and stored on infrastructure you control. orbtrace.com hosts only this marketing site. There is no telemetry beacon, no usage call-home, no 'product analytics'.

Air-gapped friendly

Runs in environments with zero outbound internet. The Causal RCA engine can call the Claude API or a fully local model (Llama 3.3, Mistral Large, vLLM) — same grounded-RAG pipeline, same citation contract.

Built for regulated industries

Fintech, healthcare, public sector, defence. Synchronous deletes, append-only audit log, GDPR Art. 17 erasure, SOC 2 Type II evidence collection, SSO/SAML, RBAC. Bring it into your compliance perimeter on day one.

Compose, Helm, or Ansible

One docker compose file for SMB teams. A Helm chart for Kubernetes shops. An Ansible playbook for bare-metal fleets. Three install paths, one product — because your platform team has opinions, and they're right.

Predictable per-cluster pricing

No per-host, per-user, per-GB, per-metric, per-anything billing. Your invoice does not change because someone added a new tag. Ever. Renewal is identical to last year, plus or minus the cluster-tier delta.

There is no Orbtrace SaaS

We will not host your telemetry. That is not a product gap — it is a product principle. We exist because the cloud option is a non-starter for our customers, and we are not going to undermine that with a hosted tier.

The killer features

From we're looking into it to here's what broke.

Causal RCA gives you the answer. Time-Travel Replay lets you test the fix against history before you ship it.

Causal RCA Engine

Hypotheses with citations, never a black box.

Anomaly detection picks the incident window. Orbtrace retrieves the relevant spans, builds the service topology, and asks Claude to reason over your data. Every conclusion cites specific span IDs — if a citation doesn't hold, the conclusion is rejected.

  • Grounded RAG over Apache Doris (your spans, your topology)
  • Citations are mandatory — no claim without a span ID
  • Embeds incidents in pgvector for similar-incident search
  • Optional fully-local model (Llama / Mistral) for air-gapped sites
See how RCA grounding works
rca.hypothesis.json
# incident: i-9f3 — checkout p99 spike
cause:     deploy svc-payments@a91c2
evidence:  312 spans · trace 7c91… · 12 deadlocks
impact:    p99 +412ms · err 0.05 → 4.2%
fix:       rollback · or release ledger.acct_lock contention
confidence 0.91 · grounded · 312 citations
Actual
deploy a91c2
spike
alert
rollback
Hypothesized
deploy a91c2 ✗
no incident
−2hincident windownow
Time-Travel Replay · ⭐ killer

Counterfactual queries on history.

Pick any moment and ask 'what if this deploy hadn't happened?' Orbtrace shows actual vs hypothesized timelines side-by-side, citing comparable past incidents and probabilistic confidence intervals.

  • Immutable span storage — replay any moment, any time
  • Built-in deploy / config-change / feature-flag timeline
  • Probabilistic outcomes with confidence intervals
  • Side-by-side actual vs counterfactual visualization
How replay reasons over history
Inside Causal RCA

Five steps from anomaly to grounded hypothesis.

The Causal RCA pipeline does not call an LLM and hope. It treats whichever model you pick (Claude, OpenAI, Azure OpenAI, Google Gemini, or a local Ollama instance) as a search-and-summarize layer over your spans — and rejects anything it can't cite back. The same engine powers the inline Explain action on any span, log line, or anomaly card.

  1. 01

    Detect

    step 1 of 5

    STL decomposition + isolation forest run continuously over Doris aggregations. The moment a service's p99, error rate, or SLO burn drifts outside the seasonal envelope, an anomaly window opens.

    pipeline · 20%
    rca.detect.ts
    anomaly.detect(
      series  = "checkout.p99",
      window  = "5m",
      baseline = "trailing 4 weeks",
    )
    → window opened: 12:34..12:39
    → anomaly_score = 0.94
  2. 02

    Retrieve

    step 2 of 5

    Orbtrace pulls the relevant spans, the service-topology graph, and recent timeline events (deploys, config changes, feature flags) for the anomaly window — straight out of Apache Doris.

    pipeline · 40%
    rca.retrieve.ts
    retrieve.context(window = anomaly.window)
    → 312 spans fetched
    → topology: 14 services, 28 edges
    → timeline:
       • 12:32 deploy svc-payments@a91c2
       • 12:30 flag.checkout.fastpath = true
  3. 03

    Reason

    step 3 of 5

    Orbtrace's AI router sends a grounded prompt to your configured provider — Claude by default, or OpenAI / Azure OpenAI / Google Gemini / a local Ollama model if you've switched in /admin/ai. The contract is strict: every claim must reference span IDs, otherwise it's rejected and the model is asked again.

    pipeline · 60%
    rca.reason.ts
    llm.complete({
      provider: env.ORBTRACE_AI_PROVIDER,  // claude | openai | azure | ollama
      system: orbtrace.rca.systemPrompt,
      context: { spans, topology, timeline },
      contract: "cite span_ids; no hallucination",
    })
    → 3 candidate hypotheses
    → best hypothesis cites 312 spans
  4. 04

    Verify

    step 4 of 5

    Citations are checked: every cited span ID must resolve in Doris and match the claim. Hypotheses without verified citations are dropped — what reaches the on-call is grounded by construction.

    pipeline · 80%
    rca.verify.ts
    verify.cite(hypothesis, citations = 312)
    → 312/312 spans resolve
    → hypothesis confidence = 0.91
    → stored in postgres + pgvector
    → similar incident lookup ready
  5. 05

    Recall

    step 5 of 5

    Verified hypotheses are embedded into pgvector and indexed alongside the incident timeline. The next time a similar pattern appears, on-call sees the past matches with their resolutions — institutional memory that survives team rotations.

    pipeline · 100%
    rca.recall.ts
    recall.similar(
      embedding = hypothesis.embedding,
      k         = 5,
      threshold = 0.82,
    )
    → 3 prior incidents matched
    → inc-2026-04-09 · resolved by deploy rollback
    → inc-2026-02-27 · resolved by ledger.acct_lock fix
    → on-call notified with playbook hints
Bring your own LLM

Pick the model. Plug in the key.

Causal RCA and Time-Travel Replay route through Orbtrace's AI router over a provider you choose at deploy time. Switch any time from /admin/ai — no migration, no vendor lock-in, no compliance review on a model your security team didn't approve.

Anthropic Claude

Default · best citation behaviour

Default provider. Strongest behaviour on the grounded-citation contract — every hypothesis comes back with span IDs that resolve.

OpenAI

Most common enterprise pick

GPT-4o family. The model your security team has probably already approved; drop in an API key and you're running.

Azure OpenAI

Compliance · data residency

The enterprise Microsoft Copilot path — same OpenAI models, Azure compliance wrappers, your tenant, your region, your audit trail.

Google Gemini

Vertex AI · long-context

The Gemini family via Vertex AI. Opt into context caching to cut grounded-RCA input-token cost ~70% on long prompts.

Ollama

Air-gapped · self-hosted

Local-only models (Llama 3.3, Mistral, Qwen — whichever you pull). The required option for air-gapped deploys; works equally well for cost-sensitive teams.

OpenAI-compatible

Self-hosted · any model

Point Orbtrace at any OpenAI-compatible server — vLLM, TGI, LocalAI, or LM Studio — with optional custom-CA or skip-verify TLS.

How it works

One active provider per deployment so RCA citations stay consistent over time. The operator picks at deploy time and hot-switches via /admin/ai; existing incident records keep their original provider attribution.

Air-gapped

No internet? Pick Ollama, pull your model of choice, point Orbtrace at the local endpoint. The entire AI surface — RCA, Replay, similar-incident embeddings — runs inside your network with zero outbound traffic.

Alerting platform

From signal to on-call, in one box.

Rules, channels, escalation, routing, schedules, silences and audit — all built in. Same data plane as RCA, so the alert that pages you links straight to the spans, traces, and topology behind it.

Seven rule families

Metric thresholds, multi-window SLO burn rate, log-volume spikes, trace-rate anomalies, anomaly-score gates, heartbeat misses, and composite (any/all) rules. Every rule is YAML-as-code or UI-edited; both round-trip.

  • metric.threshold
  • slo.burn
  • log.volume
  • trace.rate
  • anomaly
  • heartbeat
  • composite

Ten native channels

Slack, Microsoft Teams, PagerDuty, Opsgenie, Discord, Twilio SMS / Voice, Email (SMTP), Zoom and a generic JSON webhook (which also covers Mattermost and custom systems). Every channel ships a Test action so the on-call rota verifies delivery before they need it.

  • slack
  • teams
  • pagerduty
  • opsgenie
  • twilio
  • discord
  • webhook

Escalation policies

Multi-step escalation with explicit ACK tracking. If the primary on-call doesn't acknowledge in N minutes, the policy escalates to secondary, then to manager — with full audit of who saw what, when.

  • multi-step
  • ack tracking
  • fallback
  • auto-resolve

Tag-based routing

Route by service tag, severity, environment, or business hours. Same alert can page on-call out-of-hours and post to a #observability-low channel during the day. No duplicated rules.

  • service tag
  • severity
  • business-hours
  • env

Maintenance & silences

Cron-based recurring maintenance windows for planned deploys, plus ad-hoc silences with explicit owner + reason + expiry. Nothing pages during a known maintenance — and nothing gets silently muted forever.

  • cron windows
  • ad-hoc silence
  • owner + reason
  • auto-expire

Audit & analytics

Every alert event (fire, ACK, escalate, resolve, silence) is logged to PostgreSQL with append-only semantics. Built-in analytics surface noisy rules, slow ACK channels, and policies that never fire — so you tune the system instead of it tuning you.

  • append-only log
  • noisy-rule report
  • channel health
  • SIEM export
Recent alert console
delivery healthy
  • 12:34:08SLO breach · checkout-api · burn 14.2×PagerDuty · #sre-oncallACK · 47s
  • 12:11:50Anomaly score 0.91 · payments · ledger.tx p99Slack · #payments-enginvestigating
  • 11:58:02Log spike · auth.refresh.failed · +312%Slack · #authACK · 1m 12s
  • 11:42:21Heartbeat restored · ingest-worker-3Webhook · opsauto-resolved
Operations console

Every admin surface, built in.

No external admin app, no second login, no glue dashboards. Every operator workflow — from drag-drop dashboards to per-service sampling budgets to air-gapped license activation — ships with the JAR.

Dashboard editor

/dashboards

Drag-drop panel composer with templated variables, multi-pillar panels (logs · traces · metrics in one widget), Visx-rendered charts, and a starter gallery. Share, fork, or pin to the home view.

  • drag-drop
  • variables
  • multi-pillar
  • templates

Service catalog & topology

/services

Auto-discovered from span relationships — no manifest. Per-service health, dependency graph, traffic share, p50/p95/p99 baselines, and SLO burn — one click drills into spans + logs for that service.

  • auto-discovered
  • topology
  • per-service health

SLO management

/admin/slos

Pin per-service p95/p99 SLO targets. Latency baseline overrides feed the trace waterfall colouring; burn-rate alerts wire into the alerting platform automatically. Multi-window, multi-burn-rate — Google SRE workbook style.

  • per-service targets
  • burn-rate alerts
  • tier colouring

Sampling policy

/admin/sampling

Per-service monthly $ budget with real-time burn. YAML editor, hot-reload to the OTel Collector tail_sampling processor. Errors, novel paths, SLO violations and p99 spans are always kept — only well-behaved background traffic is sampled.

  • $ budgets
  • burn meter
  • YAML editor
  • hot reload

Self-monitoring

/admin/health

Orbtrace observes Orbtrace. The server emits OTLP spans, logs, and metrics into its own pipeline — every deployment, every incident gets caught by the platform itself. The strongest possible product warranty.

  • dogfooded
  • OTLP self-export
  • Prometheus scrape

License & users

/admin/license

Air-gapped activation via offline token (no internet call, ever). RBAC with role inheritance, OIDC / SAML SSO, per-tenant data partitioning, and short-lived JWT with refresh-token rotation.

  • offline activation
  • RBAC
  • OIDC
  • SAML
What you get

Eleven things no on-prem tool currently ships in one box.

Tempo gives you traces. Loki gives you logs. Prometheus gives you metrics. Orbtrace gives you all three plus the diagnosis story on top — without sending a byte to a third party, and powered by the LLM you already trust.

No Lost Trace

Adaptive buffers preserve every error and SLO-violating trace during incidents — exactly when you need them.

Continuous Context

Automatic async/queue context reconstruction across Kafka, RabbitMQ, SQS, Redis Streams, NATS. No more orphan spans.

Cost-aware Sampling

Per-service monthly budgets in real dollars. Sampling never touches the error path. Predictable spend, no custom-metric tax.

Causal RCA Engine

Grounded RAG over retrieved spans and topology. Every conclusion cites specific span IDs — no hallucination, no matter which provider you pick.

Bring Your Own LLM

Claude, OpenAI, Azure OpenAI, Google Gemini, local Ollama, or any OpenAI-compatible server — operator picks at deploy time. Hot-switch via /admin/ai. No vendor lock-in, no model your security team didn't approve.

Time-Travel Replay

Counterfactual queries on history: what if this deploy hadn't happened? Side-by-side actual vs hypothesized timelines.

Always-on Anomaly Detection

STL decomposition + isolation forest run continuously over Doris aggregations. Detects p99, error-rate, and SLO-burn drift outside the seasonal envelope — feeds RCA and the alerting platform without any model tuning.

Dynamic Latency Baselines

Each span coloured against its own service's recent p50/p95/p99 — not a one-size-fits-all 50ms/1s/5s scale. A 1.2s span is fast on a batch service, alarming on auth. Orbtrace knows the difference.

SLO-Aware Tier Colouring

Pin per-service p95/p99 contracts. The waterfall paints every span against the promise you made — green when within target, rose when over. The accountability story SREs actually want.

Auto-Discovered Service Topology

Service graph built from span relationships — no manifest, no manual map. Per-service health, dependency direction, and traffic share are first-class. Click any node to drill into spans, logs, and SLO burn for that service.

Unified Pillars

Logs, traces, metrics in a single OLAP store (Apache Doris). One query language. No more pivot-tab fatigue.

How it works

Drop in our OTel pipeline. Start diagnosing in minutes.

Orbtrace ingests through the OpenTelemetry Collector and stores in Apache Doris — the same composable pattern HyperDX, SigNoz, and Uptrace chose. The product layer is where Orbtrace earns its keep: Causal RCA, Time-Travel Replay, async stitching, cost-aware sampling.

  1. 01

    Your services

    OpenTelemetry SDK · auto-instrument

  2. 02

    OTel Collector

    Doris exporter · tail sampling · async stitching

  3. 03

    Apache Doris

    logs · traces · metrics · OLAP

  4. 04

    Orbtrace Server

    query · RCA · replay · alerting

  5. 05

    Orbtrace UI

    search · dashboards · alerts · admin

docker-compose.yml
yaml
services:
  otelcol:
    image: otel/opentelemetry-collector-contrib:latest
    command: ["--config=/etc/otel/config.yaml"]
    volumes: ["./otelcol-config.yaml:/etc/otel/config.yaml"]
    ports: ["4317:4317", "4318:4318"]
  doris-fe:
    image: apache/doris:fe-3.0.x
    environment: [DORIS_USER, DORIS_PASSWORD]
  orbtrace:
    image: nivorbit/orbtrace:0.1
    depends_on: [doris-fe, postgres, valkey]
    ports: ["8080:8080"]
    environment:
      ORBTRACE_DORIS_URL: jdbc:mysql://doris-fe:9030/orbtrace
      ORBTRACE_PG_URL:    jdbc:postgresql://postgres:5432/orbtrace
      ORBTRACE_AI_PROVIDER: anthropic
otelcol-config.yaml
yaml
receivers:
  otlp: { protocols: { grpc: { endpoint: 0.0.0.0:4317 },
                       http: { endpoint: 0.0.0.0:4318 } } }
processors:
  batch: { send_batch_size: 100000, timeout: 10s }
exporters:
  doris:
    endpoint: http://doris-fe:8030
    mysql_endpoint: doris-fe:9030
    database: orbtrace
    create_schema: false   # Orbtrace owns the schema
service:
  pipelines:
    traces:  { receivers: [otlp], processors: [batch], exporters: [doris] }
    logs:    { receivers: [otlp], processors: [batch], exporters: [doris] }
    metrics: { receivers: [otlp], processors: [batch], exporters: [doris] }
Drop-in instrumentation

Wire it up in your stack, not ours.

Orbtrace is OpenTelemetry-native. Keep your existing SDK choices — JVM, Node, Go, Python, Rust, .NET, Ruby, PHP, or any other OpenTelemetry SDK. Point the OTLP exporter at the bundled Collector and you're ingesting in seconds.

Spring Boot · Java
java
// build.gradle.kts
implementation("io.opentelemetry.instrumentation:opentelemetry-spring-boot-starter")

# application.yml
otel:
  exporter:
    otlp:
      endpoint: http://orbtrace-otelcol:4317
      protocol: grpc
  service:
    name: checkout-api
  resource:
    attributes:
      deployment.environment: prod
ASP.NET Core · .NET
csharp
// Program.cs
builder.Services.AddOpenTelemetry()
  .ConfigureResource(r => r.AddService("checkout-api"))
  .WithTracing(t => t
    .AddAspNetCoreInstrumentation()
    .AddHttpClientInstrumentation()
    .AddOtlpExporter(o =>
    {
      o.Endpoint = new Uri("http://orbtrace-otelcol:4317");
      o.Protocol = OtlpExportProtocol.Grpc;
    }));
Go
go
// main.go
import (
  "go.opentelemetry.io/otel"
  "go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc"
  "go.opentelemetry.io/otel/sdk/resource"
  sdktrace "go.opentelemetry.io/otel/sdk/trace"
)

exp, _ := otlptracegrpc.New(ctx,
  otlptracegrpc.WithEndpoint("orbtrace-otelcol:4317"),
  otlptracegrpc.WithInsecure())

tp := sdktrace.NewTracerProvider(
  sdktrace.WithBatcher(exp),
  sdktrace.WithResource(resource.Default()))

otel.SetTracerProvider(tp)
Node.js · TypeScript
ts
// instrumentation.ts
import { NodeSDK } from "@opentelemetry/sdk-node";
import { OTLPTraceExporter } from "@opentelemetry/exporter-trace-otlp-grpc";
import { getNodeAutoInstrumentations } from "@opentelemetry/auto-instrumentations-node";

new NodeSDK({
  serviceName: "checkout-api",
  traceExporter: new OTLPTraceExporter({
    url: "http://orbtrace-otelcol:4317",
  }),
  instrumentations: [getNodeAutoInstrumentations()],
}).start();
Python
py
# bootstrap.py
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter

provider = TracerProvider()
provider.add_span_processor(
  BatchSpanProcessor(
    OTLPSpanExporter(endpoint="orbtrace-otelcol:4317", insecure=True)
  )
)
trace.set_tracer_provider(provider)
Any language · OTLP/HTTP
bash
# No SDK? Any language that can POST JSON works.
curl -X POST http://orbtrace-otelcol:4318/v1/traces \
  -H 'Content-Type: application/json' \
  -d '{
    "resourceSpans": [{
      "resource": { "attributes": [
        { "key": "service.name", "value": { "stringValue": "checkout-api" } }
      ]},
      "scopeSpans": [{ "spans": [{
        "name": "POST /checkout",
        "traceId": "5b8aa5a2d2c872e8321cf37308d69df2",
        "spanId": "051581bf3cb55c13",
        "kind": 2,
        "startTimeUnixNano": "1714000000000000000",
        "endTimeUnixNano":   "1714000000231000000"
      }]}]
    }]
  }'
Kubernetes · Helm
yaml
# values.yaml — Helm
otelCollector:
  endpoint: orbtrace-otelcol.observability.svc.cluster.local:4317
sampling:
  policyEndpoint: http://orbtrace.observability.svc:8080/api/sampling/decision
  budget: { service: checkout-api, monthly: $400 }
auth:
  oidc:
    issuer: https://keycloak.your-company.internal
    audience: orbtrace
Doris query
sql
-- Find every span that touched the deadlock during the incident.
SELECT trace_id, span_id, service_name, duration_ms, status_code
FROM otel_traces
WHERE timestamp BETWEEN '2026-04-29 12:32:00' AND '2026-04-29 12:39:00'
  AND span_attributes['db.statement'] LIKE '%acct_lock%'
  AND status_code = 'ERROR'
ORDER BY duration_ms DESC
LIMIT 312;
Under the hood

Boring. Predictable. Built on standards your auditors approve of.

No proprietary protocols. No bespoke storage engine. Just well-known building blocks composed into a self-hosted product.

Ingestion is the OTel Collector's job — we don't reinvent it. Orbtrace earns its keep on what comes after: Causal RCA, Time-Travel Replay, async stitching, cost-aware sampling, alerting.

Integrations

Plays nicely with everything you already run.

No proprietary protocol. No agent rewrite. If it speaks OpenTelemetry, OIDC, SAML, or a webhook, it works with Orbtrace today — on your servers.

  • Kubernetes
  • Apache Kafka
  • PostgreSQL
  • Redis
  • PagerDuty
  • Slack
  • GitHub Actions
  • Keycloak
  • OpenLDAP
  • Anthropic Claude
  • Ollama
  • Llama 3.3
  • Mistral
  • MinIO
  • Helm
  • Terraform
  • Ansible
  • ArgoCD
  • Jenkins
  • Opsgenie
  • Microsoft Teams
  • GitLab CI
  • Kubernetes
  • Apache Kafka
  • PostgreSQL
  • Redis
  • PagerDuty
  • Slack
  • GitHub Actions
  • Keycloak
  • OpenLDAP
  • Anthropic Claude
  • Ollama
  • Llama 3.3
  • Mistral
  • MinIO
  • Helm
  • Terraform
  • Ansible
  • ArgoCD
  • Jenkins
  • Opsgenie
  • Microsoft Teams
  • GitLab CI

OpenTelemetry SDKs

  • JavaJVM
  • .NETC#
  • Gogo
  • Node.jsTS
  • Pythonpy
  • Rustrs
  • Rubyrb
  • PHPphp

Async transports (stitched)

  • Apache Kafkastream
  • RabbitMQqueue
  • AWS SQSqueue
  • Redis Streamsstream
  • NATSpubsub
  • Google PubSubpubsub

Identity & access

  • KeycloakOIDC
  • OktaSAML
  • Azure ADOIDC
  • Google WorkspaceOIDC
  • JumpCloudSAML
  • OpenLDAPLDAP

Alerting & on-call

  • PagerDutypage
  • Opsgeniepage
  • Slackchat
  • MS Teamschat
  • Mattermostchat
  • Webhook*

CI / CD events

  • GitHub Actionsci
  • GitLab CIci
  • Jenkinsci
  • ArgoCDgitops
  • Spinnakercd
  • Fluxgitops

Infrastructure

  • Kubernetesk8s
  • Nomadscheduler
  • Docker Swarmcluster
  • Terraformiac
  • Ansibleconfig
  • Helmk8s

AI inference

  • Anthropic Claudecloud / on-prem
  • Ollamalocal
  • llama.cpplocal
  • vLLMself-host
  • Llama 3.370B
  • Mistral Largeopen

Storage extensions

  • MinIOS3 cold
  • AWS S3cold
  • Cephobject
  • PostgreSQLmetadata
  • Valkeycache
  • pgvectorembeddings

Don't see what you need? The OpenTelemetry exporter, webhook alerter, and OIDC adapter cover >90% of the long tail. Webhook translators map deploy + feature-flag events from GitHub, GitLab, LaunchDarkly and a generic JSON shape onto the incident timeline — auto-correlation with no glue code. See the full integration catalogue →

Security & compliance

Built to satisfy your auditor , not just your CTO.

Self-hosting alone is not compliance. We've sat across the table from enough auditors to know what evidence they ask for — and built it in from day one.

Data sovereignty

Telemetry stays inside your perimeter — full stop. Orbtrace ships as a single docker compose, a Helm chart, or an Ansible playbook for bare-metal. No call-home, no telemetry of telemetry.

  • self-hosted
  • no telemetry beacon
  • air-gapped supported

Identity & access

OIDC and SAML 2.0 out of the box. RBAC with role inheritance and per-service data filtering — managed in /admin/users with audit on every grant. Per-tenant partitioning, short-lived JWTs with refresh-token rotation. Zero-trust by default.

  • OIDC
  • SAML
  • RBAC + inheritance
  • per-service filter
  • short-lived JWT

Audit log

Every authentication event, query, deletion, and config change is logged to PostgreSQL with append-only semantics. Export to your SIEM via webhook or Kafka.

  • append-only
  • SIEM-ready
  • tamper-evident

GDPR & the right to erase

Synchronous deletes in Apache Doris mean a GDPR Art. 17 erasure request actually erases — not eventually, but by the time the API returns 204.

  • Art. 17
  • synchronous delete
  • DSR ready

Cryptography

TLS 1.3 in transit, AES-256-GCM at rest, FIPS-140-2 mode for regulated workloads. Bring your own KMS — HashiCorp Vault, AWS KMS, GCP KMS, or PKCS#11 HSM.

  • TLS 1.3
  • AES-256-GCM
  • BYOK
  • HSM

Compliance posture

Built to support SOC 2 Type II, ISO 27001, and HIPAA evidence collection. Pre-built control mappings for your auditor — we've done the spreadsheet.

  • SOC 2 II
  • ISO 27001
  • HIPAA
Standards we're built around
GDPR
Art. 17 erasure ready
SOC 2 Type II
Pre-audit preparation
ISO 27001
Control mappings ready
HIPAA
On-prem data residency
FIPS 140-2
Encryption standards
PCI-DSS 4.0Roadmap
On the roadmap
FedRAMPRoadmap
On the roadmap (air-gapped)
BSI C5Roadmap
On the roadmap

Want our security packet (architecture diagrams, threat model, control mappings)? Email security@nivorbit.com →

Honest comparison

No marketing fluff. Here's where we differ.

We use the same OpenTelemetry primitives as everyone else. The difference is what we build on top — and the fact that we will never ask you to ship telemetry off-prem.

  • Self-hosted, on-prem only

    Yes
    Datadog
    No
    HyperDX
    Partial
    Tempo / Jaeger
    Partial
    Uptrace / SigNoz
    Partial
  • Air-gapped install (no internet)

    Yes
    Datadog
    No
    HyperDX
    Partial
    Tempo / Jaeger
    Yes
    Uptrace / SigNoz
    Yes
  • Unified logs + traces + metrics

    Yes
    Datadog
    Yes
    HyperDX
    Yes
    Tempo / Jaeger
    No
    Uptrace / SigNoz
    Yes
  • Causal RCA with grounded LLM

    Yes
    Datadog
    Yes
    HyperDX
    Partial
    Tempo / Jaeger
    No
    Uptrace / SigNoz
    Partial
  • Time-Travel Replay

    Yes
    Datadog
    No
    HyperDX
    No
    Tempo / Jaeger
    No
    Uptrace / SigNoz
    No
  • Continuous anomaly detection (STL + isolation forest)

    Yes
    Datadog
    Yes
    HyperDX
    No
    Tempo / Jaeger
    Partial
    Uptrace / SigNoz
    Partial
  • Async / queue context stitching

    Yes
    Datadog
    Yes
    HyperDX
    Partial
    Tempo / Jaeger
    No
    Uptrace / SigNoz
    Yes
  • Per-service dynamic latency baselines

    Yes
    Datadog
    Yes
    HyperDX
    No
    Tempo / Jaeger
    No
    Uptrace / SigNoz
    Partial
  • SLO-aware span tier colouring

    Yes
    Datadog
    Partial
    HyperDX
    No
    Tempo / Jaeger
    No
    Uptrace / SigNoz
    No
  • Cost-aware $-budget sampling

    Yes
    Datadog
    Partial
    HyperDX
    No
    Tempo / Jaeger
    No
    Uptrace / SigNoz
    No
  • Built-in alerting (rules · channels · escalation · routing)

    Yes
    Datadog
    Yes
    HyperDX
    Partial
    Tempo / Jaeger
    No
    Uptrace / SigNoz
    Yes
  • Built-in dashboard editor

    Yes
    Datadog
    Yes
    HyperDX
    Yes
    Tempo / Jaeger
    No
    Uptrace / SigNoz
    Yes
  • No custom-metric tax

    Yes
    Datadog
    No
    HyperDX
    Yes
    Tempo / Jaeger
    Yes
    Uptrace / SigNoz
    Yes
  • Predictable, perpetual pricing

    Yes
    Datadog
    No
    HyperDX
    Partial
    Tempo / Jaeger
    Yes
    Uptrace / SigNoz
    Yes
  • Compose · Helm · Ansible install paths

    Yes
    Datadog
    No
    HyperDX
    Yes
    Tempo / Jaeger
    Yes
    Uptrace / SigNoz
    Yes
  • Runs on an 8 GB VPS

    Yes
    Datadog
    No
    HyperDX
    Yes
    Tempo / Jaeger
    Yes
    Uptrace / SigNoz
    Partial

Comparison reflects publicly available capabilities as of June 2026. We are not affiliated with these projects.

What teams say

On-prem teams finally have a tool that tells them what broke.

Datadog wanted us to ship customer PII through a public-cloud SaaS. That was a non-starter for our compliance team. Orbtrace runs in our DMZ and the RCA quality has significantly improved our mean-time-to-understanding.
Platform engineering lead
EU fintech · 600 services
Improved
MTTU
We replaced a Tempo + Loki + Grafana stack and the on-call experience changed overnight. The incident page links to relevant spans and citations. We're not going back.
SRE manager
Healthcare SaaS · regulated workloads
Increased
stability
Time-Travel Replay let me prove to leadership that the rollback was the right call without re-running the broken deploy. That was the moment I sold the team on Orbtrace.
Staff engineer
Public-sector contractor
Actionable
post-mortem

Quotes are from design partners under NDA. Logos available on request.

Pricing

Quoted per deployment. No metering.

Commercial on-prem software with a single annual licence — quoted per deployment so procurement gets one clean number, not a calculator. No tiers, no per-user / per-host / per-GB billing — capacity scales with your hardware, not your contract. Optional services (pilot, hourly support, implementation) are quoted separately so you only pay for what you actually use.

Annual licence

Orbtrace Licence

Customquote· annual

Single annual fee. Unlimited Doris BE nodes, unlimited services, unlimited users. Renewal years are maintenance + support only.

  • Causal RCA + Time-Travel Replay
  • Async Stitching + cost-aware sampling
  • Bring-your-own LLM (Claude / OpenAI / Azure / Gemini / Ollama)
  • SSO / SAML / OIDC + RBAC + audit log
  • Unlimited hosts, services, retention
  • Air-gapped activation, no phone-home

Optional services

30-day Pilot

Full product in your environment, real workloads, sales-led onboarding. Convert or walk away.

PriceFree
Premium Support

Slack channel, 8h response SLA, troubleshooting + custom dashboard help. Pay only for hours actually used — no retainer.

Price$200-300 / hr
Implementation

2-4 week fixed-bid engagement — cluster sizing, hardening, observability standup, custom dashboard portfolio.

Pricefrom $15K
Typical SaaS APMVariablemetered monthly bill
Orbtrace On-premPredictableannual/perpetual licence

Shift from metered SaaS to predictable on-premise observability — your data stays in your network.

Sales-led. No download portal, no shared installer. Every deployment goes through onboarding so the cluster is sized, hardened, and observable from the first hour.

Frequently asked

Honest answers to the questions every SRE team asks.

  • Is Orbtrace really on-premise only?

    Yes. There is no Orbtrace SaaS — we deploy as a single docker compose (or Helm chart, or Ansible playbook) inside your network. Our cloud only hosts the marketing site. Your telemetry never leaves your infrastructure.

  • Why Apache Doris instead of ClickHouse?

    Doris speaks the MySQL wire protocol — zero learning curve for Java teams and any tool that talks SQL. It also ships native inverted indexes that beat ClickHouse on full-text log search by 3–10× in our benchmarks, has synchronous deletes (important for GDPR), and simpler cluster management.

  • How does Time-Travel Replay actually work?

    Doris stores every span with full attribute history — immutable. Replay builds a timeline of deploys, config changes, and feature-flag toggles, then asks the RCA engine to reason about a counterfactual using past comparable incidents. The output is a probabilistic timeline with confidence intervals, never a deterministic claim.

  • What does Orbtrace cost in production?

    Sales-led, on-prem only. One annual licence covers unlimited hosts, services, and users — no tiers, no per-user / per-host / per-GB metering anywhere. Pricing is quoted per deployment so customers with strict procurement or compliance constraints get a clean number rather than a calculator. Optional services (30-day pilot, $200-300/hr premium support, $15K+ fixed-bid implementation projects) are quoted separately so you only pay for what you actually consume.

  • How does Orbtrace integrate with PagerDuty / Opsgenie / Slack?

    Native senders for all three, plus Microsoft Teams, Discord, Zoom, Twilio SMS/Voice and email — and a generic signed JSON webhook that covers Mattermost or any custom system. Alert rules are defined as code (YAML / Terraform provider) and stored in PostgreSQL. Multi-window, multi-burn-rate SLO alerts are first-class — same model as Google's SRE workbook.

  • How do I silence alerts during planned maintenance?

    Two primitives. Maintenance Windows (cron-based) suppress matching alerts on a recurring schedule — perfect for nightly batch jobs or weekly deploys. Ad-hoc Silences are one-off, with required owner + reason + auto-expiry so nothing gets muted forever. Both are managed in /admin/alerts and audited like every other policy change.

  • How do you know Orbtrace works in production?

    We dogfood it. Every Orbtrace server emits OTLP spans, logs, and metrics into its own pipeline, and our engineering team's incident response uses this exact product. /admin/health surfaces the self-monitoring view so operators can verify the platform is observing itself before they trust it with their workloads.

  • Does the AI / RCA need internet access?

    Only if you pick a hosted provider. Causal RCA defaults to Anthropic Claude, but at deploy time you can switch to OpenAI, Azure OpenAI, Google Gemini, or a local Ollama endpoint via /admin/ai. Air-gapped customers run Ollama (Llama 3.3, Mistral, Qwen — whichever they pull) with zero outbound traffic. The grounded-RAG pipeline is identical across providers — citations to span IDs are mandatory either way.

  • Will I have to rip out my OpenTelemetry instrumentation?

    No — Orbtrace is OpenTelemetry-native. Point your existing OTLP exporters at the bundled OTel Collector and you're done. We do not require any proprietary SDK or agent.

  • How does cost-aware sampling avoid dropping the bug I'm chasing?

    Sampling never touches the error path or anomalies. The policy endpoint Orbtrace exposes to the OTel Collector's tail_sampling processor enforces hard rules: errors, novel paths, SLO violations, and p99-latency spans are always kept. Only well-behaved background traffic is sampled when budget pressure is high.

  • What hardware do I need?

    Plan for 8 GB RAM as the realistic floor — the default Compose stack runs three Doris storage nodes, so 4 GB only works if you trim Doris to a single node. Production starts around 2 vCPU / 8 GB RAM for the orbtrace-server JAR plus a single-node Doris (4 vCPU / 16 GB); for real production traffic plan for 16–32 GB RAM and SSD storage. A 3-node Doris cluster comfortably ingests 100K+ events/sec on modest hardware.

  • Can I run Causal RCA fully offline (no hosted LLM)?

    Yes. Set ORBTRACE_AI_PROVIDER=ollama and point Orbtrace at a local Ollama endpoint. We've benchmarked Llama 3.3 70B and Mistral Large for grounded-RAG quality and ship recommended prompts for both. Connected deploys can equally well stay on Anthropic Claude (default), OpenAI, Azure OpenAI, or Google Gemini — same /admin/ai surface, same RCA contract.

  • How does air-gapped license activation work?

    Upload an offline license token at /admin/license. The server validates the signature locally and unlocks features — no internet call, no phone-home, ever. Renewal works the same: a fresh token from sales, dropped into the same screen. The only network egress Orbtrace makes is the OTLP / SMTP / webhook traffic you explicitly configure.

Stop staring at dashboards. Start fixing incidents.

Pilot Orbtrace inside your network for 30 days. Bring real telemetry, watch Causal RCA hypotheses land during the first incident.

hello@nivorbit.com

On-prem only. No telemetry beacon, no usage call-home. Your data never leaves your network — including during the pilot.