Sentrinel β€” Feature Matrix & Roadmap

Researched July 2026 against: Sentry, PostHog, Axiom, Better Stack, Highlight.io (now LaunchDarkly Observability), Langfuse, Helicone (maintenance mode since the Mintlify acquisition), Grafana + Prometheus, Datadog / New Relic, and the OpenTelemetry spec.

Two market signals worth remembering when scoping: Highlight.io was acquired and folded into LaunchDarkly, and Helicone is in maintenance mode β€” being an inline proxy or chasing every category is risky; a focused API-monitoring product with a shared ingest pipeline is the durable shape.

Capability matrix (βœ“ first-class Β· ◐ partial Β· – absent)

Capability Sentry PostHog Axiom Better Stack Langfuse Grafana+Prom Datadog/NR Sentrinel today Sentrinel target
Multi-tenant orgs + auth ◐ βœ“ βœ“ βœ“ βœ“ ◐ βœ“ βœ“ βœ“
Endpoint/APM metrics (p50/p95/p99, Apdex) βœ“ – ◐ ◐ ◐ ◐ βœ“ βœ“ βœ“
Error grouping & stack traces βœ“ βœ“ – ◐ – – βœ“ βœ“ fingerprinted issues βœ“
Request logs w/ headers & payloads ◐ – βœ“ βœ“ – βœ“ βœ“ βœ“ βœ“
Application logs correlated to requests ◐ – βœ“ βœ“ – βœ“ βœ“ βœ“ βœ“
Log search & live tail ◐ – βœ“ βœ“ – βœ“ βœ“ βœ“ search + SSE tail βœ“
Field/attribute explorer (log vocabulary) – – βœ“ ◐ – – ◐ βœ“ keys + top values, click to filter βœ“
Breadcrumbs (trail leading to a failure) βœ“ – – – – – βœ“ βœ“ logs+spans+outcome interleaved βœ“
Issue assignment & ownership βœ“ – – ◐ – – βœ“ βœ“ assign + filter by assignee βœ“ + ownership rules
Distributed tracing (OTLP ingest) βœ“ – βœ“ ◐ βœ“ βœ“ βœ“ βœ“ nested waterfall + OTLP βœ“
Uptime / synthetic checks ◐ – – βœ“ – ◐ βœ“ βœ“ (HTTP checks) βœ“ multi-region later
Cron / heartbeat monitoring βœ“ – – βœ“ – ◐ ◐ βœ“ βœ“
Alerting + notification channels βœ“ ◐ βœ“ βœ“ ◐ βœ“ βœ“ βœ“ Slack, Discord, webhook, email βœ“
SLOs & error budgets – – ◐ ◐ – βœ“ βœ“ βœ“ βœ“
Consumers / API-client analytics – ◐ ◐ – ◐ – ◐ βœ“ βœ“
Host metrics (CPU/RAM) – – ◐ βœ“ – βœ“ βœ“ βœ“ βœ“
Database query performance (per-query metrics) – – – ◐ – ◐ βœ“ βœ“ pg_stat_statements, ranked by share of DB time βœ“ + MySQL
Database wait events & blocking chains – – – – – ◐ βœ“ βœ“ 1s pg_stat_activity sampling, pg_blocking_pids βœ“
Database health (connections, cache hit, deadlocks, lag) – – – ◐ – βœ“ βœ“ βœ“ βœ“
Table bloat & index advisories – – – – – – ◐ βœ“ bloat, missing index, index overhead βœ“ + unused-index detection
Optional query-text masking in the collector – – – – – – ◐ βœ“ off by default, literals stripped on the host when on βœ“
Explain plans & plan-change history – – – – – – βœ“ βœ“ collector EXPLAINs the top statements, one row per plan change βœ“
Custom dashboards (saved charts) βœ“ βœ“ βœ“ βœ“ βœ“ βœ“ βœ“ βœ“ composed in Settings, rendered at /dashboards βœ“
Release / deploy tracking βœ“ ◐ – – – ◐ βœ“ βœ“ deploy markers βœ“
LLM call tracing & cost ◐ βœ“ – – βœ“ – βœ“ βœ“ priced GenAI spans βœ“
SQL access to telemetry ◐ βœ“ βœ“ βœ“ ◐ βœ“ βœ“ ◐ guarded, Postgres-backed βœ“ ClickHouse-backed
Status pages – – – βœ“ – – ◐ βœ“ public + incidents βœ“
Custom / business metrics – ◐ βœ“ ◐ – βœ“ βœ“ βœ“ counter, gauge, histogram; folded in-process, in TypeScript, Python and Dart βœ“
Session replay βœ“ βœ“ – – – – βœ“ βœ“ rrweb, error-triggered, masked by default βœ“ + sampled sessions
Mobileβ†’backend on one trace ◐ – – – – ◐ ◐ βœ“ traceparent from the SDK βœ“ β€” the thing nobody else joins cleanly
Product analytics (events, funnels, retention) – βœ“ – – – – ◐ βœ“ web + mobile track(), funnels, retention βœ“
Feature flags – βœ“ – – – – ◐ – out of scope
Profiling βœ“ – – – – βœ“ βœ“ – out of scope for now
On-call schedules / escalation ◐ – – βœ“ – βœ“ βœ“ – integrate, don't build
Saved queries / starred views βœ“ βœ“ βœ“ ◐ – βœ“ βœ“ βœ“ named filter sets per page, starrable βœ“
Query language (APL / Discover) ◐ βœ“ βœ“ ◐ – βœ“ βœ“ ◐ guarded SQL console, against ClickHouse or Postgres field filters cover most of it
Release health (crash-free rate per release) βœ“ – – – – – ◐ βœ“ crash-free sessions + users, per release βœ“
Mobile crash reporting βœ“ – – – – – ◐ βœ“ uncaught Dart, persisted across the crash ◐ native iOS/Android absent
Mobile performance (app start, frames) βœ“ – – – – – ◐ βœ“ start-to-first-frame, slow/frozen frames βœ“
Breadcrumbs on a crash report βœ“ – – – – – βœ“ βœ“ 25 bounded, rendered as a trail βœ“
Browser error monitoring βœ“ ◐ – – – βœ“ βœ“ βœ“ uncaught + rejections + fetch, no key in the bundle βœ“
Browser β†’ backend distributed trace βœ“ – – – – βœ“ βœ“ βœ“ traceparent on same-origin fetch βœ“
Request geo + client IP βœ“ ◐ ◐ βœ“ – ◐ βœ“ βœ“ country + resolved client IP per request βœ“
Serving host per request ◐ – βœ“ βœ“ – βœ“ βœ“ βœ“ which vhost/replica answered βœ“
Dark mode βœ“ βœ“ βœ“ βœ“ βœ“ βœ“ βœ“ βœ“ light / dark / auto βœ“
Trials, quotas & billing enforcement βœ“ βœ“ βœ“ βœ“ ◐ – βœ“ βœ“ trial clock, blocking gate, ingest grace βœ“ + self-serve checkout
Operator console (per-tenant billing) ◐ ◐ ◐ ◐ – – βœ“ βœ“ allowlisted platform admin βœ“
Web release health (crash-free page loads) βœ“ – – – – – ◐ βœ“ one session per page load βœ“
Source maps / minified stack recovery βœ“ – – – – βœ“ βœ“ βœ“ upload by API key from CI; symbolicated at ingest, before fingerprinting βœ“
Web vitals (LCP, CLS, INP) βœ“ βœ“ – – – βœ“ βœ“ βœ“ collected per session, shown as p75 per release βœ“
Stack symbolication / deobfuscation βœ“ – – – – – βœ“ ◐ dSYM/ProGuard artifacts upload and store; no native stack is rewritten yet βœ“
Suspect commits / code owners βœ“ – – – – – ◐ – needs VCS integration
Purpose-bound API keys (one key per integration) ◐ ◐ βœ“ ◐ ◐ – βœ“ βœ“ server, mobile, database collector, OTLP, AI agent β€” enforced per surface, a leak exposes one integration βœ“
Coding-agent access (MCP server + CLI) ◐ – ◐ – – – ◐ βœ“ six read tools + resolve, on an agent key pinned to one app βœ“
Durable ingest with two-way failover – – βœ“ ◐ – ◐ βœ“ βœ“ log used whenever a broker answers; store↔log failover behind a circuit breaker; state on /health βœ“
Python / Django SDK βœ“ ◐ ◐ ◐ ◐ ◐ βœ“ βœ“ requests, errors, logs, consumers, custom metrics, spans, outbound tracing, browser tunnel, resource usage; no runtime deps βœ“
Python / FastAPI & Starlette SDK βœ“ ◐ ◐ ◐ ◐ ◐ βœ“ βœ“ pure ASGI middleware: requests, errors (HTTPException and validation included), logs from sync and async endpoints, spans, httpx tracing, browser tunnel; no runtime deps βœ“

Priorities

P0 β€” core (shared pipeline, natural fit)

  1. βœ… Done β€” Storage architecture for scale: Postgres control plane + ClickHouse telemetry store.
  2. βœ… Done β€” Error grouping into fingerprinted issues (issues table, /api/issues, Issues page with resolve/ignore/regression lifecycle).
  3. βœ… Done β€” Application logs: console capture in @sentrinel/plugin correlated by request ID, Logs tab + log-level filter in Request logs UI.
  4. βœ… Done β€” Alert notification channels: Slack / Discord / generic webhook delivery, managed on the Alerts page. (Email still open.)
  5. βœ… Done β€” OTLP/HTTP ingest (POST /v1/traces) so any OTel SDK or Collector can send traces.
  6. βœ… Done β€” Auth & multi-tenancy: self-serve signup (one org per account), sessions, roles, team management, login throttling, and org scoping enforced on every dashboard route (SENTRINEL_REQUIRE_AUTH=true).
  7. βœ… Done β€” Quota enforcement: per-org monthly counters, 429 over quota, usage on Settings.

P1 β€” cheap extensions of P0 8. βœ… Done β€” Cron/heartbeat monitoring: check-in URLs, sweeper, missed-job notifications. 9. βœ… Done β€” Live tail (SSE) on requests, logs, and errors, with a Live toggle on Request logs. 10. βœ… Done β€” SLOs: availability + latency objectives with error budgets and burn-down. 11. βœ… Done β€” Custom dashboards: saved widget configs (Settings β†’ Custom dashboards). 12. βœ… Done β€” Deploy markers: version from the plugin β†’ deployments, shown on Settings.

P2 β€” with clear pull only 13. βœ… Done β€” Status pages: public route over uptime + incidents (/public/status/:slug). 14. βœ… Done β€” LLM spans: GenAI attrs + price table β†’ cost per model, client, and endpoint. 15. βœ… Done β€” Guarded read-only SQL console.

P3 β€” mobile 16. βœ… Done β€” Crash reporting for Dart: uncaught errors captured automatically and written to disk before the process dies, delivered on the next launch. The persistence is what makes it crash reporting rather than error reporting β€” the ordinary buffer flushes on a timer a crashing app never reaches. 17. βœ… Done β€” Release health: one app_sessions row per launch, crash-free sessions and crash-free users per release. Force-quits are counted as abnormal, apart from crashes, because blaming a release for a user swiping the app away makes the number worthless. 18. βœ… Done β€” Flutter integration package (sentrinel_flutter): framework and engine error handlers, app start, slow/frozen frames, navigation breadcrumbs. Separate from the pure-Dart core so the core stays usable from CLIs and server-side jobs.

P4 β€” the gaps the first list named 19. βœ… Done β€” Email notification channel, refused at creation on a deployment that cannot actually send, rather than failing silently at the moment an alert fires. 20. βœ… Done β€” SQL console against ClickHouse, not just Postgres, so the console can see the store that holds the telemetry. 21. βœ… Done β€” Saved views: a named filter set per page, starrable, scoped to the org. 22. βœ… Done β€” Web vitals: LCP, CLS and INP through PerformanceObserver, carried on the session and reported as p75 per release β€” vitals have a long tail and the tail is the point, so an average would hide it. 23. βœ… Done β€” Explain plans and plan-change history: the collector EXPLAINs the top statements on the host and sends one row per plan change, so a regression shows as the plan that changed rather than a query that got slow. 24. βœ… Done β€” Source maps: uploaded with an API key from CI (a build pipeline has no browser session, which is why the feature was unreachable while the only upload route needed one), symbolicated at ingest before fingerprinting so issues group by real source location, with the minified stack preserved.

Not started, and deliberately so β€” native iOS/Android crash handlers and native stack symbolication (source maps, the web half, shipped). They are one project, not two: native traces without symbolication are unreadable hex, so shipping the first without the second gives you nothing. Together that is a quarter of work plus permanent maintenance across both toolchains, on Sentry's strongest ground. The differentiator worth defending instead is already shipped: a mobile tap and the backend span it caused on one trace, via traceparent.

Closed since this list was written: the email notification channel, the ClickHouse-backed SQL console, saved views, web vitals, explain plans and plan-change history, and the source-map half of symbolication. Each is marked in the matrix above with what actually shipped.

Remaining, in the order the gap costs you something:

# Gap Where it stands
1 Native stack symbolication symbol_artifacts accepts and stores dsym, proguard and symbols, and lib/symbol-service.ts reads only sourcemap. So the upload half is done and nothing consumes the native half β€” a stored dSYM changes no stack today.
2 Native iOS/Android crash handlers Deliberately not started β€” see above. Pointless before #1 regardless.

Smaller, already scoped in the matrix above: unused-index detection, MySQL support, multi-region uptime checks, sampled (not just error-triggered) session replay, and issue ownership rules.

P3 β€” deliberately out of scope RUM, product analytics, feature flags, surveys (browser SDK products); profiling (cost/benefit); LLM gateway/proxying (different risk posture β€” see Helicone); on-call scheduling (integrate PagerDuty/Better Stack via the webhook channel instead).