
Sutradhar
sutradhar - the one who holds the thread and directs the play
Agents that cut the alert noise, find the cause, and fix what is safe to fix.
Sutradhar is an agentic IT operations console: AI agents that watch a company's IT systems, catch problems, work out what is wrong, and fix them automatically where it is safe, asking a person to approve anything risky. It covers the whole loop, from cutting noisy alerts down to real incidents, through diagnosing root cause and weighing risk before acting, to executing or proposing fixes and reporting results to leadership. The approval boundary is explicit and every action is recorded.
Workflow
Problem
IT operations drowns in alerts that mostly are not incidents, so real problems are diagnosed late and the same routine fixes are performed by hand every week.
Outcome
Noise collapsed into real incidents, root cause diagnosed, and safe remediation executed with anything risky escalated for approval.
01
Compress the noise
Alert volume is reduced to the set of incidents actually occurring.
02
Diagnose the cause
Each incident is traced across services, infrastructure, and recent changes to a probable cause.
03
Weigh the risk
Candidate remediations are assessed for blast radius before anything is executed.
04
Act or escalate
Safe fixes run automatically; anything consequential waits for a human, and both paths are recorded.
Modules
Alert Compression
Noisy alert volume reduced to the incidents that are actually happening.
Root-Cause Diagnosis
Incidents traced to cause across services, infrastructure, and recent change.
Risk-Weighted Action
Each candidate fix weighed for blast radius before anything executes.
Approval & Reporting
Safe fixes executed, risky ones escalated, and outcomes reported to leadership.
Representative use cases
Collapsing Alert Storms into the Incidents That Are Actually Happening
An enterprise operations centre receives tens of thousands of monitoring alerts a week across infrastructure, applications, and network.
Diagnosing Root Cause Across Services, Infrastructure, and Recent Change
A platform team's mean time to diagnosis is dominated by the search for what changed, across deployments, configuration, and infrastructure.
Executing Safe Remediations While Escalating Anything with Blast Radius
An operations team performs the same routine remediations weekly - restarts, cache clears, capacity adjustments - by hand, at all hours.
Reporting Operational Health to Leadership Without a Manual Pack
An IT leadership team receives a monthly operations pack assembled by hand from six systems, arriving three weeks after the period it describes.
Bringing Agentic Operations Inside a Regulated Change-Control Process
A financial services operations team wants agentic remediation but works under a change-control regime that requires every production change to be recorded and attributable.
Evidence
What Wayam can show for this entry
Curated solution design with an owner and a review date. No performance or production claim.
Allowed claims at this tier: Possible workflow, typical stack, prerequisites.
- Evidence tier
- T4 · Reference pattern
- Integration status
- Typical enterprise system
- Metric status
- Scenario only
- Evidence owner
- Not yet attached
- Validated
- Not yet attached
- Valid until
- —
- Content version
- 2026.09
Pending evidence attachment
This entry represents an enterprise architectural reference pattern. Customer benchmarks, run replays, and metric verifications are established during technical discovery.
Limitations
- Headline metrics are representative until a customer result is attached.
Representative metrics · Reference · scenario only
Noise to incidents
alerts compressed before a human sees them
Approval-gated
anything with real blast radius
These describe the intended outcome of the design. They are not measured customer results until a validated case is attached above.
Architecture & controls
Composed across the operating loop
Ingest
Connect the systems of record and read the signals the work already produces.
Covered by the platform
Reason
Ground context, score options against policy and the stated goal, draft the next step.
4 agent roles
Act
Execute an approved step in the system of record and keep the evidence.
3 agent roles
Govern
Set policy, gate consequential actions on a named approver, and audit what ran.
3 agent roles
Pack Architecture & Agent Composition Graph
10 Composed Agents- AGT-0222
Incident Triage Agent
IT Operations
- AGT-0223
Root Cause Analysis Agent
IT Operations
- AGT-0224
Change Risk Assessment Agent
IT Operations
- AGT-0278
Natural Language to SQL Agent
Data and Analytics
- AGT-0279
Data Catalog Curation Agent
Data and Analytics
- AGT-0280
Data Quality Monitoring Agent
Data and Analytics
- AGT-0345
Quality Inspection Reporting Agent
Quality
- AGT-0348
Scrap and Rework Analysis Agent
Quality
Control gates
- A named approver on every consequential write-back
- Evidence and audit trail kept with every run
Capability atoms
- signal-ingestion
- anomaly-detection
- root-cause-reasoning
- alert-routing
- autonomous-execution
- workflow-orchestration
- human-approval
- evidence-audit-trail
Typical enterprise systems
Typical stack for solution design; compatibility is validated during discovery.
- ServiceNow ITSM
- Datadog
- PagerDuty
- Kubernetes
- Snowflake
- dbt
- Monte Carlo
- Power BI
Designed for
technology, financial-services, public-sector
Business case
Model a Sutradhar scenario with your own baseline
Three scenarios from the numbers you enter. Capacity released is time; it becomes a saving only when roles or costs are actually removed or avoided.
| Scenario | Improvement | Capacity released | Cashable savings | Net annual | Payback |
|---|---|---|---|---|---|
| Conservative | 9% | — | — | — | — |
| Expected | 18% | — | — | — | — |
| Upside | 24% | — | — | — | — |
Enter an annual volume and a baseline cost to see figures.
Illustrative planning scenario, not a guarantee. Results depend on process baseline, adoption, data quality, integration scope, controls, and deployment costs.
Pilot this
From solution design to a bounded pilot
Step 1 · 3–4 weeks
Discovery Sprint
Validate the workflow, data, controls, baseline and business case before anything is built.
Gate: Blueprint and pilot plan signed by the business owner, technical owner and Wayam.
Step 2 · 6–10 weeks
Bounded Pilot
Prove quality and value on agreed data against agreed acceptance tests.
Gate: Acceptance thresholds met on the evaluation set; go/no-go decision recorded.
Pairs well with