TB-Scale per Hour
Peak production Trace ingestion
The Production Data Pipeline for Agent Traces
Capture model responses, tool calls, execution spans, and application events in one S3-backed lakehouse. Turn high-volume, evolving JSON into query-ready and eval-ready data—without operating a fragmented Trace pipeline.
VALIDATED BY A LEADING FOUNDATION MODEL COMPANY
Peak production Trace ingestion
Cumulative Trace data volume
Spans in a single Trace
Maximum size of a single Trace
Online Trace data can reach TBs per hour and grow to trillions of records, pushing traditional databases beyond sustained ingestion limits.
A single long-running task can invoke thousands of tools, process millions of context tokens, and generate complex, multi-level spans.
Every model or tool upgrade can introduce new JSON fields, forcing frequent pipeline changes when ingestion depends on a fixed schema.
Teams often stitch together Kafka, Flink, Airflow, a data warehouse, and object storage—creating longer pipelines and higher operational costs.
One Data Foundation from Agent Execution to Model Improvement
Databend Cloud unifies raw Trace storage, incremental SQL transformation, scheduling, and analytics in one managed pipeline.
Preserve the complete execution context in S3, process only new events, and continuously produce query-ready Trace models for Evals, debugging, replay, attribution, and training.
Scroll horizontally to explore the full diagram.
Sustain TB-scale Trace ingestion per hour with parallel loading and horizontal worker scaling. Keep writes stable as agent traffic and Trace volume grow.
Preserve evolving, deeply nested JSON, then extract and search the fields that matter—without rebuilding the pipeline for every schema change.
Use Stream and Task for in-database processing and scheduling, without separate Flink or Airflow pipelines. Independent warehouses isolate workloads, while the managed service shortens the path to production.
Store complete raw traces in customer-controlled lakehouse, with full control over retention periods, field extraction, and analytical models for each agent and application.
| Capability | What It Delivers |
|---|---|
| Native VARIANT | Preserves deeply nested and evolving JSON without defining every field upfront. |
| Stream + Task | Processes newly arrived events with incremental SQL pipelines. |
| Cluster Key + Recluster | Reduces scanning when retrieving long-running Traces by time and trace_id. |
| Elastic Warehouses | Isolates ingestion, transformation, querying, and maintenance workloads. |
| S3-Backed Storage | Supports customer-controlled retention and historical reprocessing. |
| PrivateLink + Masking | Protects sensitive prompts, model outputs, and application data. |
Reconstruct model calls, tool invocations, decisions, and parallel branches by trace_id.
Transform production Trace data into datasets for regression testing and model comparison.
Reproduce failures and validate changes to models, prompts, tools, and agent harnesses.
Connect quality, cost, and latency to the decisions and components that produced them.
Convert successful execution paths into structured assets for post-training and reinforcement learning.
Build a production-ready Agent Trace pipeline on AWS—from raw execution events to debugging, evaluation, replay, attribution, and training.