Behavioral trajectory simulator
Know what happens next, before it happens
Shadow forecasts what your live user segments will do if you change nothing — the decline, the confidence band, and the worst case — grounded in your own product data and backtested for calibration. Built for CTOs, product, and engineering leaders.
Current → p50
63% → 0%
Decline prob
0%
Fewer activations · 30d
~0
How to read it: the line is the median outcome (p50). The cone is the range 80% of simulated futures fall inside — its floor is your worst case. It widens with time because uncertainty compounds the further out you forecast.
Not synthetic personas
Shadow runs on your real users' behavior — not made-up profiles or interviewed strangers.
Not another churn score
Full 7/14/30-day trajectories with calibrated confidence bands and a worst case — not a single opaque number.
Not AI testing
Shadow forecasts your live users, not your code or click flows. It answers "what happens if we do nothing?"
Analytics explains what happened.
Shadow shows what happens next.
Every insight includes a forward-looking trajectory. If Shadow cannot simulate with confidence, it shows INSUFFICIENT SIGNAL.
Segment-Level Forecasts
7/14/30-day trajectories for every simulated segment, with a p10–p90 confidence band: 80% of simulated futures land inside it, and p10 is your realistic worst case.
Cost of Inaction
Every forecast quantifies what doing nothing costs — e.g. reviews lost over 30 days — with risk and escalation probability.
Production Safe
Read-only connector. No writes, no PII, no production impact. Encrypted and isolated by design.
Fast Integration
Connect via Amplitude or your data warehouse. See insights in 1–3 days, no code changes required.
Backtested & calibrated — not just a confident number
In the Yelp demo, actual behavior landed inside Shadow's confidence band 84% of the time (target 80%), at 1.6pp mean error — scored across 7, 14, 21 and 30-day horizons on generated fixtures.
How a number gets made
Every figure Shadow shows you is traceable to one of three engines — and you can always see which one produced it. The math is deterministic. Only the words are AI.
Deterministic Path Simulation
Models system behavior: feature flags, eligibility rules, screen navigation. Answers "what paths are possible or blocked?"
Statistical Cohort Simulation
Uses historical distributions, regression, and Monte Carlo to forecast cohort behavior. Primary engine for Shadow.
AI Interpretation
Reads the numbers L1 and L2 produced and explains them: plausible drivers, what's uncertain, and what would prove the projection wrong. It never computes or overrides a number. Fires only when you ask "why?"
Three things Shadow talks about — don't mix them up
L1–L3 is how a number is produced. Modes are how your app is executed to observe it — modelled as a flow graph, or replayed in a real client. And Preview → What-If → owned models is what questions you can ask, over time. Every answer in the product carries its layer and mode, so you always know which engine you're trusting.
From intro call to first forecast in 3–5 working days
A Shadow forward-deployed engineer does the work. Your side of it is roughly four hours of people-time, spread across the week — no code changes, no SDK, no data migration.
Scoping call
Day 0 · 45 minYou: One product leader and one data or analytics person. You name the metric that worries you and the cohorts you care about.
Shadow: We confirm your event taxonomy can support it and tell you honestly if it can't yet.
Read-only access
Day 1 · 60 minYou: Issue an Amplitude Export API key (or a read-only warehouse role). Your infra or data team, one ticket.
Shadow: We deploy the ShadowConnector — a Docker container that runs inside your infrastructure, pulls read-only, and never writes back. No PII leaves your boundary.
Ingest and event mapping
Days 1–2 · asyncYou: Answer questions over Slack as they come up. Usually 30 minutes total.
Shadow: We pull 90 days of history, infer your metric definitions, and map raw events to the behaviour they represent. This is the step that needs a human — event taxonomies are always messier than the docs say.
Segment discovery and gating
Days 2–3 · asyncYou: Nothing.
Shadow: We enumerate candidate cohorts, run data-sufficiency checks on each, and discard the ones too thin to forecast honestly. You will see exactly which segments were excluded and why.
Backtest and calibration
Days 3–4 · asyncYou: Nothing.
Shadow: We re-run the model at past cutoffs against what actually happened on your data, and publish the coverage. If the bands are not calibrated on your product, you find out before you trust a single forecast.
Readout
Day 5 · 60 minYou: Bring whoever owns the roadmap.
Shadow: We walk your ranked risk segments, the cost of inaction, and the calibration evidence. You leave with a shortlist of cohorts to act on — or a clear statement that nothing is currently at risk, which is also a valid answer.
What can slow this down. Fewer than 28 days of clean history, event volumes below the sufficiency thresholds, or an event taxonomy where the same user action fires under three different names. We will tell you in the scoping call if we think you're in that position — a forecast built on insufficient signal is worse than no forecast, so we will not ship one.
Read-only. Encrypted. Isolated.
Shadow is a read-only simulation layer that fits into the same trust boundary as Amplitude, Mixpanel, or your observability pipeline — with stricter isolation. It never takes action on users or production systems.
Shadow never takes action on users or production systems
No writes to production. No outbound user communication. No side effects (emails, billing, webhooks). No autonomous decisions.
Data Minimization
No raw PII ingested by default. User identities are represented by anonymized or hashed identifiers. Only the minimum behavioral state required to simulate is used.
- Aggregated behavioral signals only
- Sensitive fields excluded, tokenized, or hashed at source
- Permission-scoped state access
Encryption & Isolation
Every layer of the stack is encrypted and sandboxed. Simulation environments are logically and cryptographically isolated from your production systems.
- TLS 1.2+ in transit, AES-256 at rest
- Dedicated sandboxed simulation environments
- No shared execution context with production
Production Separation
Shadow traffic is explicitly tagged and separated. Simulations run outside your runtime with no access to write paths, user-facing channels, or billing systems.
- Read-only connector, never writes back
- Tagged and separated from production traffic
- Security team approved in a single review
Priced like infrastructure, not seats
Usage-based on the users and segments you monitor — not per head. Read-only and no-PII, inside the same trust boundary as your analytics.
You already spend six figures a year on analytics and observability — Amplitude, Datadog, Pendo. Shadow is the forward-looking layer on top, at a fraction of that.
Starter
Single team, one data source.
$1,000/mo
Up to 250k monitored MAU
- Segment forecasts + confidence bands
- Backtesting & calibration
- Amplitude or Demo connector
- On-demand AI interpretation
- Email support
Every plan includes the read-only connector, calibrated confidence + backtesting, and the INSUFFICIENT SIGNAL guardrail. Indicative pricing — design-partner terms for the first cohort.