Skip to main content

Product Suite

Flashduty is a unified observability platform designed for DevOps, SRE, and operations teams, providing end-to-end solutions from monitoring and alerting to incident response.

AI SRE

Autonomous troubleshooting Agent

On-call

Accelerate alert response

RUM

Real user experience monitoring

Monitors

Unified monitoring platform

AI SRE Autonomous Troubleshooting

A conversational, autonomous SRE Agent: issue instructions in natural language, and the AI investigates incidents, finds root causes, calls tools, and accumulates reusable operational knowledge. Integrated deeply with Flashduty incident response and IM collaboration.
  • Chat to troubleshoot: the Agent plans, calls tools, and streams its investigation and conclusion
  • Integrated with incident response: spin up a session from an incident or war room with full context
  • Knowledge that compounds: DUTY.md-rooted Knowledge Packs hold long-lived operational context
  • Extensible tooling: Skills, MCP, A2A Agents, and self-hosted Runners
  • Start an investigation from the console or an IM group, visible to the whole team
  • Automatic initial diagnosis posted back to the incident war room
  • Use /insight to review the last 30 days and quantify operational friction
  • Build reusable runbooks and troubleshooting flows
AI SRE is available to accounts on an On-call Pro or higher subscription, with no application needed. Starting September 16, 2026, 08:00 Beijing time, it’s billed on actual usage. See the AI SRE introduction and billing terms.

Product overview

Learn the full capability map and console navigation

Get started

Create a session and start conversational troubleshooting

On-call Alert Management

A unified intelligent alert response platform: reduce noise, schedule, assign, escalate, and notify to help teams respond to and handle production incidents quickly.
  • Intelligent noise reduction: Alert grouping, inhibition, and deduplication to reduce 90% of alert noise
  • Flexible assignment: Multi-level escalation, dynamic routing, and rotation scheduling
  • Multi-channel notifications: Feishu/Lark, Dingtalk, WeCom, Slack, phone calls, SMS
  • Rich integrations: Native support for 50+ alert sources
  • Build 7x24 on-call systems to ensure service availability
  • Aggregate alerts from multiple sources into a unified incident management portal
  • Establish escalation mechanisms to ensure timely response
  • Analyze alert data to continuously improve monitoring quality

Quick Start

Complete your first alert integration in 5 minutes

Product Comparison

In-depth comparisons with PagerDuty and Opsgenie

RUM User Experience Monitoring

Real User Monitoring helps you understand how real users experience your application and quickly identify and resolve frontend issues.
  • Performance monitoring: Full-chain tracking of page loads, resource loads, and API calls
  • Error tracking: Automatic collection and grouping of JS errors and network errors
  • Session replay: Reproduce user operation paths to quickly replicate issues
  • Custom metrics: Report business-specific metrics to meet personalized needs
  • Monitor core web vitals (LCP, FID, CLS) for web applications
  • Track and analyze frontend errors to improve application stability
  • Analyze user behavior paths to optimize product experience
  • Correlate with backend traces for full-stack observability

Quick Start

Integrate SDK to start monitoring

Performance Analysis

Learn about performance metrics and analysis methods

Monitors Management

Onboard data sources scattered across network zones and manage alert rules centrally; rules are evaluated locally by the alert engine deployed next to your data, and the events they produce flow into On-call.
  • Multiple data source types: Prometheus, VictoriaLogs, Loki, Elasticsearch, SLS, and Tencent Cloud CLS for metrics and logs, plus MySQL, PostgreSQL, Oracle, ClickHouse, Redis, Kafka, and MongoDB for databases and middleware
  • Local evaluation next to your data: monitedge runs inside your private network, syncs rules from SaaS and queries data sources locally, so data never leaves your network; engine instances in the same cluster shard the rules automatically
  • Three evaluation modes: threshold, no-data, and any-data, with consecutive or cumulative hit counting and per-level recovery semantics
  • Unified management views: manage rules in a folder tree, with overview, active alerts, Query Workbench, Entity Tree and dashboards, plus rule import/export and change audit
  • Monitoring is spread across multiple clouds and data centers, and you want to onboard sources and manage alert rules in one place
  • Data sources live in a private network without public access, and alert evaluation should stay local
  • Manage Prometheus, log, database, and middleware alerting under one workflow
  • Let rule-generated alerts flow straight into Flashduty grouping, escalation, and routing

Quick Start

Create your first monitoring task

FAQ

Answers to common questions

Developers

Integrate Flashduty through Open API and Webhooks for automation and custom development.

Quick Start

Authentication, request specs, error handling

API Catalog

All 354 endpoints organized by module

About Pagination

Traditional and cursor pagination

Contact Us

Technical Support

Scan to add our WeCom for one-on-one technical supportTechnical Support WeCom

Business Inquiries

Scan to add our business manager on WeComBusiness Manager WeCom

Console Feedback

Sign in to the console and submit feedback in the bottom left corner