Product Suite
Flashduty is a unified observability platform designed for DevOps, SRE, and operations teams, providing end-to-end solutions from monitoring and alerting to incident response.AI SRE
Autonomous troubleshooting Agent
On-call
Accelerate alert response
RUM
Real user experience monitoring
Monitors
Unified monitoring platform
AI SRE Autonomous Troubleshooting
A conversational, autonomous SRE Agent: issue instructions in natural language, and the AI investigates incidents, finds root causes, calls tools, and accumulates reusable operational knowledge. Integrated deeply with Flashduty incident response and IM collaboration.Core capabilities
Core capabilities
- Chat to troubleshoot: the Agent plans, calls tools, and streams its investigation and conclusion
- Integrated with incident response: spin up a session from an incident or war room with full context
- Knowledge that compounds: DUTY.md-rooted Knowledge Packs hold long-lived operational context
- Extensible tooling: Skills, MCP, A2A Agents, and self-hosted Runners
Use cases
Use cases
- Start an investigation from the console or an IM group, visible to the whole team
- Automatic initial diagnosis posted back to the incident war room
- Use /insight to review the last 30 days and quantify operational friction
- Build reusable runbooks and troubleshooting flows
AI SRE is available to accounts on an On-call Pro or higher subscription, with no application needed. Starting September 16, 2026, 08:00 Beijing time, it’s billed on actual usage. See the AI SRE introduction and billing terms.
Product overview
Learn the full capability map and console navigation
Get started
Create a session and start conversational troubleshooting
On-call Alert Management
A unified intelligent alert response platform: reduce noise, schedule, assign, escalate, and notify to help teams respond to and handle production incidents quickly.Core Capabilities
Core Capabilities
- Intelligent noise reduction: Alert grouping, inhibition, and deduplication to reduce 90% of alert noise
- Flexible assignment: Multi-level escalation, dynamic routing, and rotation scheduling
- Multi-channel notifications: Feishu/Lark, Dingtalk, WeCom, Slack, phone calls, SMS
- Rich integrations: Native support for 50+ alert sources
Use Cases
Use Cases
- Build 7x24 on-call systems to ensure service availability
- Aggregate alerts from multiple sources into a unified incident management portal
- Establish escalation mechanisms to ensure timely response
- Analyze alert data to continuously improve monitoring quality
Quick Start
Complete your first alert integration in 5 minutes
Product Comparison
In-depth comparisons with PagerDuty and Opsgenie
RUM User Experience Monitoring
Real User Monitoring helps you understand how real users experience your application and quickly identify and resolve frontend issues.Core Capabilities
Core Capabilities
- Performance monitoring: Full-chain tracking of page loads, resource loads, and API calls
- Error tracking: Automatic collection and grouping of JS errors and network errors
- Session replay: Reproduce user operation paths to quickly replicate issues
- Custom metrics: Report business-specific metrics to meet personalized needs
Use Cases
Use Cases
- Monitor core web vitals (LCP, FID, CLS) for web applications
- Track and analyze frontend errors to improve application stability
- Analyze user behavior paths to optimize product experience
- Correlate with backend traces for full-stack observability
Quick Start
Integrate SDK to start monitoring
Performance Analysis
Learn about performance metrics and analysis methods
Monitors Management
Onboard data sources scattered across network zones and manage alert rules centrally; rules are evaluated locally by the alert engine deployed next to your data, and the events they produce flow into On-call.Core Capabilities
Core Capabilities
- Multiple data source types: Prometheus, VictoriaLogs, Loki, Elasticsearch, SLS, and Tencent Cloud CLS for metrics and logs, plus MySQL, PostgreSQL, Oracle, ClickHouse, Redis, Kafka, and MongoDB for databases and middleware
- Local evaluation next to your data:
monitedgeruns inside your private network, syncs rules from SaaS and queries data sources locally, so data never leaves your network; engine instances in the same cluster shard the rules automatically - Three evaluation modes: threshold, no-data, and any-data, with consecutive or cumulative hit counting and per-level recovery semantics
- Unified management views: manage rules in a folder tree, with overview, active alerts, Query Workbench, Entity Tree and dashboards, plus rule import/export and change audit
Use Cases
Use Cases
- Monitoring is spread across multiple clouds and data centers, and you want to onboard sources and manage alert rules in one place
- Data sources live in a private network without public access, and alert evaluation should stay local
- Manage Prometheus, log, database, and middleware alerting under one workflow
- Let rule-generated alerts flow straight into Flashduty grouping, escalation, and routing
Quick Start
Create your first monitoring task
FAQ
Answers to common questions
Developers
Integrate Flashduty through Open API and Webhooks for automation and custom development.Quick Start
Authentication, request specs, error handling
API Catalog
All 354 endpoints organized by module
About Pagination
Traditional and cursor pagination
Contact Us
Technical Support
Scan to add our WeCom for one-on-one technical support

Business Inquiries
Scan to add our business manager on WeCom

Console Feedback
Sign in to the console and submit feedback in the bottom left corner