DrDroid
Backed by Y Y Combinator

Self-learningAI SRE Agent.

DrDroid connects to your cloud, code, and telemetry, builds a live knowledge graph of your stack, and gives SRE engineers faster root-cause analysis and automated remediation, cutting MTTR on every incident.

Map your entire stack in one graph.

DrDroid maps every repo, dashboard, K8s pod and cloud resource into one live knowledge graph — then traces the blast radius across connected entities the moment an alert fires.

payments-api SERVICE Repo 3 payments-api Dashboard 7 p95-latency Alerts 124 p95 anomaly Issues 18 GH#4821 Resources 9 k8s pods Infra 14 aws rds
Infra dependency confirmed 3 pods healthy 14 correlations 1 anomaly 2 deploys today
01

Cross-tool correlation

A GitHub repo maps to a Datadog service, a Grafana dashboard, K8s pods, and AWS resources, automatically.

02

Decision engine

When an alert fires, the graph traces the blast radius across every connected entity in seconds.

03

Continuously learning

Every alert, deploy, and incident strengthens the graph, surfacing patterns no dashboard can.

How DrDroid learns your stack.

We connect to your existing tools, crawl all telemetry, and generate a knowledge graph of your stack.

01 / CONNECT

Read-only access to your entire stack.

OAuth into cloud, code, CI/CD, and observability. No agents. No code changes. Live in 30 minutes.

AWS · GCP · Azure GitHub · GitLab Datadog · Grafana · NR
02 / CRAWL

We crawl all telemetry and build your knowledge graph.

Metrics, logs, traces, cloud configs, repos, docs, runbooks. All crawled and mapped into a cross-tool knowledge graph. Which repo → which service → which dashboard → which pods. Always live, always learning.

knowledge graph context map always live
03 / ACT

Act with full context.

The knowledge graph powers proactive suggestions, root-cause diagnosis, and automated runbooks — all with full context.

proactive explainable guarded
Tighten retry budget · orders-svc SUGGEST
Cause of INC-4821 · sidecar OOM RCA · 9m
Auto-scale on memory pressure RUN
Drain node-12 · disk-full RUN
Self-learning agent

Lower MTTR with every investigation.

Every alert, deploy, and incident strengthens the graph. The second time the same pattern shows up, DrDroid already knows the answer — so MTTR keeps dropping.

First encounter 34s
01get_metrics("latency.p95")not found
02retry: trace.duration.p95spike found
03check recent deploysdead end
04check config changesdead end
05check pod metricsOOM found
Same alert, next day 12s
01get_metrics("trace.duration.p95")remembered
02get_pod_metrics()OOM confirmed
6 tool calls, 2 dead ends, 1 error 3 tool calls, 0 dead ends 65% faster
Learn more about how the agent learns

One memory for your entire stack.

AI Memory holds your service graph, runbooks, docs, and every live signal, alerts, deploys, conversations, incidents. It builds patterns over time so every engineer starts with full context, not a blank slate.

drdroid.app / ai-memory
⌘K

Memory Explorer

Platform Knowledge
memory
Metric/ 22,103
Panels/ 2,208
Daily logs/ 1,375
Infrastructure Components/ 689
Dashboards/ 646
Services/ 622
Runbooks/ 59
Communication/ 33
Repo context/ 7
Alert Rules/ 4
Skills/ 4
MCP Assets/ 2
Alerts & Activity
alerts
Alerts/ 6,198
Issues/ 1,501
Recent Changes/ 1,467
Investigations/ 224
Human Conversations/ 66

Classified Alerts View

1 Hour 4 Hours 24 Hours Custom
Relevant Alerts 134
infra APITimeoutError on OpenAI API in podracer 2 alerts
Last: a few minutes ago Sentry
infra APITimeoutError on Azure cognitive services endpoint 2 alerts
Last: a few minutes ago sentry
code psycopg2 UndefinedColumn created_at protoproddb connector 2 alerts
Last: a few minutes ago sentry
code psycopg2 UndefinedColumn tool_calls protoproddb connector 1 alert
Last: a few minutes ago sentry
code PostgreSQL UndefinedColumn investigation_id protoproddb 1 alert
Last: a few minutes ago sentry
Suppressed Alerts 46
known-noise 46 alerts
Last seen: 9 minutes ago sentry +3 more reports +2 more

Service Catalog

Service Name Upstream Downstream Data Sources Created By Rule Source
azure_monitorinfra None None 3 sources DroidAgentV2 Rules managed
app_serviceservice None None 3 sources DroidAgentV2 Rules managed
addon-resizerinfra None None 9 sources DroidAgentV2 Rules managed
storageinfra None None 9 sources DroidAgentV2 Rules managed
network_watcherinfra None None 9 sources DroidAgentV2 Rules managed
metrics-serverinfra None None 14 sources DroidAgentV2 Rules managed

Connect your entire stack.

Cloud, code, observability, incident response and ticketing — wired in via integrations, read-only and reversible.

Cloud & Infra
AWS AWS
Google Cloud Google Cloud
Azure Azure
Kubernetes Kubernetes
Amazon EKS Amazon EKS
GKE GKE
Code & Delivery
GitHub GitHub
GitHub Actions GitHub Actions
Bitbucket Bitbucket
Jenkins Jenkins
Argo CD Argo CD
Observability
Datadog Datadog
Grafana Grafana
New Relic New Relic
Prometheus Prometheus
Elastic Elastic
SignOz SignOz
Incident & Response
PagerDuty PagerDuty
OpsGenie OpsGenie
Sentry Sentry
Rootly Rootly
Zenduty Zenduty
Rollbar Rollbar
Workflow & Ticketing
Slack Slack
MS Teams MS Teams
Linear Linear
Jira Jira
Notion Notion
Confluence Confluence

See DrDroid in action

Watch how engineering teams use DrDroid to cut MTTR and stay ahead of incidents.

Built for SRE engineers on call.

We measure ourselves on pages avoided and minutes saved during the incident, not dashboards rendered.

"Earlier, debugging meant hopping between logs, workflows, and infra dashboards trying to piece together what went wrong. DrDroid pulls the context together and points us in the right direction, even someone new to the system can figure things out."

Rahul Bhattacharya Rahul Bhattacharya · Co-founder & CTO, Adopt.ai

"One time I was woken up at 3am by a pager that escalated. I instantly asked DrDroid to investigate it and in a few minutes, I was able to close the issue directly from Slack."

Moiz Arsiwala Moiz Arsiwala · CTO, WorkIndia

"DrDroid understood our context too well. It gave recommendations which showed deep understanding of the infrastructure and helped reduce 20–30% cost."

Prateek Prateek · Head of Technology, Stanza Living

Enterprise-ready security and deployment.

DrDroid runs where your data lives, meets the bar your security team sets, and ties its pricing to outcomes you actually care about.

SOC2 SOC 2 Type II certified Read-only integrations SSO / SAML
01

Self-hosted deployment

Run entirely inside your VPC or on-prem. No data leaves your network. Deploy via Helm or Docker Compose with air-gapped support.

02

Outcome guarantees

We tie our success to yours — measurable reduction in MTTR and incident frequency, SLA-backed with quarterly reviews.

03

Security & compliance

SOC 2 Type II, encrypted at rest and in transit, read-only access to all integrations. Built to pass your vendor review on day one.

Generate your knowledge graph in minutes.

Connect your stacks and see your services mapped in minutes.