Ruby SDK for Langfuse — the open-source LLM engineering platform. Tracing, prompt management, and evaluation for Ruby LLM apps, with first-class support for the Langfuse v4 observations-first data model.
New projects should use ingestion_mode: :otel (Langfuse v4). That path sends traces over OTLP/HTTP (/api/public/otel/v1/traces) with x-langfuse-ingestion-version: 4, so data shows up in real time and observation-level evaluators, cost, and the Observations API v2 work as designed. The tracing API (Langfuse.trace, #generation, #span, #agent, …) is unchanged; only the transport and ID format differ.
On Langfuse Cloud, POST /api/public/ingestion stops accepting everything except scores on 16 November 2026. Self-hosted v4 should use OTEL as well. See the Langfuse v4 guide and the official custom-ingestion migration.
- ⚡ Langfuse v4 / OpenTelemetry: OTLP ingestion, W3C hex IDs,
usage_details/cost_details, observation types (agent,tool,retriever, …) - 🔍 Tracing: Traces, spans, generations, events, and typed observations
- 📝 Prompt Management: Versioned prompts with a bounded cache and stale-on-outage reads
- 📊 Evaluation: Built-in evaluators and scores (trace, observation, session, dataset run)
- 🚀 Async Processing: Background batching, queue bounds, fork-safe flush,
at_exitshutdown - 🔒 Resilience: Typed errors, Retry-After + jittered backoff, null-object degradation for
Langfuse.trace
This gem requires Ruby >= 3.1 and is tested against Ruby 3.1–4.0. For
development, Ruby version is managed with mise (defaults
to the latest stable Ruby via .mise.toml).
Add this line to your application's Gemfile:
gem 'langfuse-ruby'And then execute:
$ bundle installOr install it yourself as:
$ gem install langfuse-ruby# Install mise (if not already installed), then trust the project config and
# install the pinned Ruby version.
brew install mise # macOS; see mise docs for other OSes
mise install # installs Ruby from .mise.toml
bundle install # install gem dependencies
bundle exec rake spec # run the test suiteThe SDK default host is US Cloud (https://us.cloud.langfuse.com). Override
host / LANGFUSE_HOST / LANGFUSE_BASE_URL for EU, Japan, HIPAA, or
self-hosted. ingestion_mode still defaults to :legacy for compatibility;
set it to :otel for v4.
require "langfuse"
Langfuse.configure do |config|
config.public_key = ENV.fetch("LANGFUSE_PUBLIC_KEY")
config.secret_key = ENV.fetch("LANGFUSE_SECRET_KEY")
config.host = ENV["LANGFUSE_HOST"] || ENV["LANGFUSE_BASE_URL"] || "https://us.cloud.langfuse.com"
config.ingestion_mode = :otel # Langfuse v4 — required for real-time OTEL ingestion
end
# Equivalent:
# Langfuse.new(..., ingestion_mode: :otel)
# LANGFUSE_INGESTION_MODE=otel| Region | host |
|---|---|
| US (SDK default) | https://us.cloud.langfuse.com |
| EU | https://cloud.langfuse.com |
| Japan | https://jp.cloud.langfuse.com |
| HIPAA | https://hipaa.cloud.langfuse.com |
| Self-hosted | your Langfuse origin (no trailing path) |
No extra gems are required. The SDK maps traces/spans/generations to OTLP JSON
and sets x-langfuse-ingestion-version: 4 on the OTEL connection.
v4 is observations-first: a trace is the set of observations that share a
trace_id. Put the overall request/response on the trace (root span) and
on the generation/span that actually produced them. Prefer usage_details /
cost_details over the legacy usage hash — v4 uses usage_details for cost.
Langfuse.trace("chat-completion", user_id: "user-123", session_id: "sess-456",
input: { message: "Hello, world!" }) do |trace|
generation = trace.generation(
name: "openai-completion",
model: "gpt-4o",
input: [{ role: "user", content: "Hello, world!" }],
model_parameters: { temperature: 0.7, max_tokens: 100 }
)
response = call_llm(...) # your code
generation.end(
output: response.content,
usage_details: { input: 12, output: 18, total: 30 }, # v4 cost model
cost_details: { input: 0.0001, output: 0.0006, total: 0.0007 }
)
generation.score(name: "faithfulness", value: 0.9)
trace.update(output: response.content)
end # flush is automatic in the block formScores always go through /api/public/ingestion (score-create), even in
:otel mode. Trace/observation IDs are normalized to W3C hex so they attach to
the OTEL-ingested spans. If an OTEL export fails mid-batch, both the OTEL
events and that batch's scores are re-queued.
Typed observations (agent, tool, chain, retriever, embedding,
evaluator, guardrail) are spans with langfuse.observation.type set. They
filter and evaluate correctly in the v4 UI.
Langfuse.trace("document-qa", user_id: "user-456", input: { query: "What is Ruby?" }) do |trace|
agent = trace.agent(name: "qa-agent", input: { query: "What is Ruby?" })
retrieval = agent.retriever(name: "vector-search", input: { query: "What is Ruby?", top_k: 5 })
retrieval.end(output: { documents: ["Ruby is a programming language..."] })
gen = agent.generation(
name: "openai-completion",
model: "gpt-4o",
input: [{ role: "user", content: "What is Ruby?" }],
prompt: Langfuse.get_prompt("qa-system") # optional prompt link
)
gen.end(
output: "Ruby is a dynamic programming language.",
usage_details: { input: 50, output: 20, total: 70 }
)
agent.end(output: { answer: "Ruby is a dynamic programming language." })
trace.update(output: { answer: "Ruby is a dynamic programming language." })
end:otel (v4, recommended) |
:legacy |
|
|---|---|---|
| Transport | OTLP/HTTP JSON /api/public/otel/v1/traces |
batched POST /api/public/ingestion |
| Header | x-langfuse-ingestion-version: 4 |
none |
| IDs | W3C 32-char trace / 16-char span hex | UUIDs |
| Create + update | collapsed into one span (v4 is append-only) | sent as separate events |
| Usage | usage normalized into langfuse.observation.usage_details + gen_ai.usage.* |
legacy usage object |
| Scores | still the ingestion API, IDs coerced to hex | ingestion API |
| Tracing API | identical | identical |
Do not dual-send the same IDs through both modes into one project. Switch with
ingestion_mode: / LANGFUSE_INGESTION_MODE (values are downcased; a typo
falls back to :legacy with a warning). Full attribute mapping, evaluator
notes, and a cutover checklist: docs/V4.md.
For most use cases, you can use the simplified class-level API with automatic flush:
require 'langfuse'
# Configure once — v4 / OTEL is the recommended ingestion path
Langfuse.configure do |config|
config.public_key = ENV["LANGFUSE_PUBLIC_KEY"]
config.secret_key = ENV["LANGFUSE_SECRET_KEY"]
config.ingestion_mode = :otel
end
# Use block-based tracing - flush happens automatically!
Langfuse.trace("my-trace", user_id: "user-1", input: { message: "Hello" }) do |trace|
generation = trace.generation(
name: "openai-chat",
model: "gpt-4",
input: [{ role: "user", content: "Hello" }],
model_parameters: { temperature: 0.7 }
)
# Call your LLM
response = call_openai(...)
# Record the response
generation.end(
output: response.content,
usage_details: { input: 10, output: 15, total: 25 } # or usage: response.usage
)
trace.update(output: response.content)
end # Automatic flush here!# Get and compile a prompt in one call
prompt = Langfuse.get_prompt("greeting-prompt", variables: { name: "Alice" })
# => "Hello Alice! How can I help you today?"
# Get prompt without compilation
prompt_obj = Langfuse.get_prompt("greeting-prompt")
compiled = prompt_obj.compile(name: "Bob")The simplified API includes null objects that ensure your code continues working even if Langfuse is unavailable:
# If Langfuse fails, a NullTrace is used - your code still runs
Langfuse.trace("my-trace") do |trace|
# This works even if Langfuse is down
gen = trace.generation(name: "test", model: "gpt-4", input: "hello")
gen.end(output: "world")
end# Get the process-wide, thread-safe singleton client
client = Langfuse.client
# Manual flush (when not using block-based tracing)
Langfuse.flush
# Shutdown the client (idempotent; also runs via at_exit when shutdown_on_exit is true)
Langfuse.shutdown
# Reset the singleton (useful for testing; prefer Langfuse.new for isolated clients)
Langfuse.reset!trace = client.trace(name: "document-qa")
# Create a span for document retrieval
retrieval_span = trace.span(
name: "document-retrieval",
input: { query: "What is machine learning?" }
)
# Add a generation for embedding
embedding_gen = retrieval_span.generation(
name: "embedding-generation",
model: "text-embedding-ada-002",
input: "What is machine learning?",
output: [0.1, 0.2, 0.3], # embedding vector
usage: { prompt_tokens: 5, total_tokens: 5 }
)
# End the retrieval span
retrieval_span.end(
output: { documents: ["ML is...", "Machine learning involves..."] }
)
# Create a span for answer generation
answer_span = trace.span(
name: "answer-generation",
input: {
query: "What is machine learning?",
context: ["ML is...", "Machine learning involves..."]
}
)
# Add LLM generation
llm_gen = answer_span.generation(
name: "openai-completion",
model: "gpt-3.5-turbo",
input: [
{ role: "system", content: "Answer based on context" },
{ role: "user", content: "What is machine learning?" }
]
)
llm_gen.end(
output: { answer: "Machine learning is a subset of AI..." },
usage_details: { input: 50, output: 30, total: 80 }
)
answer_span.end(output: { answer: "Machine learning is a subset of AI..." })Langfuse v4 queries observations directly. Use a specific type so traces
filter and evaluate correctly. Helpers exist on Client, Trace, Span, and
Generation. as_type: on #span does the same thing; an explicit as_type:
passed to a typed helper cannot override that helper's type.
| Helper | langfuse.observation.type |
Use for |
|---|---|---|
#span |
span |
generic timed work |
#generation |
generation |
LLM calls (model, tokens, cost, prompt link) |
#event |
event |
point-in-time logs |
#agent |
agent |
orchestration / tool-calling loops |
#tool |
tool |
a single function or API call |
#chain |
chain |
stitching retrieval → generation, etc. |
#retriever |
retriever |
vector store / DB lookups |
#embedding |
embedding |
embedding model calls (model / usage go into metadata) |
#evaluator / #evaluator_obs |
evaluator |
scoring functions (Client#evaluator is an alias of #evaluator_obs) |
#guardrail |
guardrail |
safety / moderation |
trace = client.trace(name: "support-agent", user_id: "u1")
agent = trace.agent(name: "planner", input: { question: "Reset my password" })
tool = agent.tool(name: "lookup-user", input: { email: "a@example.com" })
tool.end(output: { user_id: "u1" })
guard = agent.guardrail(name: "content-filter", input: { text: "Reset my password" })
guard.end(output: { blocked: false })
gen = agent.generation(name: "reply", model: "gpt-4o", input: [...])
gen.end(output: "I can help with that.", usage_details: { input: 40, output: 12, total: 52 })
agent.end(output: { reply: "I can help with that." })Create generic events for custom application events and logging:
# Create events from trace
event = trace.event(
name: "user_action",
input: { action: "login", user_id: "123" },
output: { success: true },
metadata: { ip: "192.168.1.1" }
)
# Create events from spans or generations
validation_event = span.event(
name: "validation_check",
input: { rules: ["required", "format"] },
output: { valid: true, warnings: [] }
)
# Direct event creation
event = client.event(
trace_id: trace.id,
name: "audit_log",
input: { operation: "data_export" },
output: { status: "completed" },
level: "INFO"
)# Get a prompt
prompt = client.get_prompt("chat-prompt", version: 1)
# Prompt names with special characters are automatically URL-encoded
prompt = client.get_prompt("EXEMPLE/my-prompt") # Works correctly!
# Compile prompt with variables
compiled = prompt.compile(
user_name: "Alice",
topic: "machine learning"
)
puts compiled
# Output: "Hello Alice! Let's discuss machine learning today."Note: Prompt names containing special characters (like
/, spaces,?, etc.) are automatically URL-encoded. You don't need to manually encode them.
Fetched prompts are cached per client for cache_ttl_seconds (default 60). The
cache is bounded (200 entries), TTLs use a monotonic clock, and if a refetch
fails while an expired entry exists, the stale copy is served with a warning —
a Langfuse outage does not break prompt resolution for prompts seen before.
# Create a text prompt
text_prompt = client.create_prompt(
name: "greeting-prompt",
prompt: "Hello {{user_name}}! How can I help you with {{topic}} today?",
labels: ["greeting", "customer-service"],
config: { temperature: 0.7 }
)
# Create a chat prompt
chat_prompt = client.create_prompt(
name: "chat-prompt",
prompt: [
{ role: "system", content: "You are a helpful assistant specialized in {{domain}}." },
{ role: "user", content: "{{user_message}}" }
],
labels: ["chat", "assistant"]
)# Create prompt templates for reuse
template = Langfuse::PromptTemplate.from_template(
"Translate the following text to {{language}}: {{text}}"
)
translated = template.format(
language: "Spanish",
text: "Hello, world!"
)
# Chat prompt templates
chat_template = Langfuse::ChatPromptTemplate.from_messages([
{ role: "system", content: "You are a {{role}} assistant." },
{ role: "user", content: "{{user_input}}" }
])
messages = chat_template.format(
role: "helpful",
user_input: "What is Ruby?"
)# Exact match evaluator
exact_match = Langfuse::Evaluators::ExactMatchEvaluator.new
result = exact_match.evaluate(
input: "What is 2+2?",
output: "4",
expected: "4"
)
# => { name: "exact_match", value: 1, comment: "Exact match" }
# Similarity evaluator
similarity = Langfuse::Evaluators::SimilarityEvaluator.new
result = similarity.evaluate(
input: "What is AI?",
output: "Artificial Intelligence is...",
expected: "AI is artificial intelligence..."
)
# => { name: "similarity", value: 0.85, comment: "Similarity: 85%" }
# Length evaluator
length = Langfuse::Evaluators::LengthEvaluator.new(min_length: 10, max_length: 100)
result = length.evaluate(
input: "Explain AI",
output: "AI is a field of computer science that focuses on creating intelligent machines."
)
# => { name: "length", value: 1, comment: "Length 80 within range" }# Add scores to traces or observations
trace = client.trace(name: "qa-session")
# Score the entire trace
trace.score(
name: "user-satisfaction",
value: 0.9,
comment: "User was very satisfied"
)
# Score specific generations
generation = trace.generation(
name: "answer-generation",
model: "gpt-3.5-turbo",
output: { content: "The answer is 42." }
)
generation.score(
name: "accuracy",
value: 0.8,
comment: "Mostly accurate answer"
)
generation.score(
name: "helpfulness",
value: 0.95,
comment: "Very helpful response"
)# Tag all events with a tracing environment (also via LANGFUSE_TRACING_ENVIRONMENT)
client = Langfuse.new(
public_key: "pk-lf-...",
secret_key: "sk-lf-...",
environment: "production"
)
# Sample a fraction of traces deterministically (also via LANGFUSE_SAMPLE_RATE)
# All events of a trace share the same keep/drop decision.
sampled_client = Langfuse.new(
public_key: "pk-lf-...",
secret_key: "sk-lf-...",
sample_rate: 0.1
)
# Mask sensitive fields before sending (applied to input/output/metadata)
masked_client = Langfuse.new(
public_key: "pk-lf-...",
secret_key: "sk-lf-...",
mask: ->(value) { value.to_s.gsub(/\b\d{16}\b/, "***CARD***") }
)Events are flushed in the background every flush_interval seconds, or as soon
as flush_at events are queued (default 15, env LANGFUSE_FLUSH_AT). Batches
are automatically split to respect the 3.5 MB ingestion API limit, and a
process-wide at_exit hook flushes pending events on shutdown.
The queue is bounded by max_queue_size (default 10,000, env
LANGFUSE_MAX_QUEUE_SIZE): while Langfuse is unreachable the oldest events are
dropped with a warning instead of growing memory without bound, and batches
that fail with a permanent error (4xx) are dropped rather than retried forever.
shutdown lets the flush thread finish its current send before returning, and
after a fork (e.g. Puma workers) each process recreates its own flush thread.
update and end send only the fields you changed, so ending a generation does
not re-upload its prompt, input or model parameters.
client = Langfuse.new(
public_key: "pk-lf-...",
secret_key: "sk-lf-...",
flush_at: 50, # flush after 50 events
flush_interval: 10, # or every 10 seconds
max_queue_size: 20_000, # drop the oldest events beyond this many queued
shutdown_on_exit: true # flush on process exit (default)
)Scores can target a trace, an observation, a session, or a dataset run, and carry metadata, a config reference, and an annotation queue link:
# Trace-level score
trace.score(name: "accuracy", value: 0.9, comment: "good")
# Observation-level score (trace_id is set automatically on Span/Generation)
generation.score(name: "faithfulness", value: 0.8, data_type: "NUMERIC")
# Session-level score
client.score(name: "csat", value: 5, session_id: "sess-1", data_type: "NUMERIC")
# Dataset-run score with metadata and config link
client.score(
name: "hallucination",
value: 0.2,
dataset_run_id: "run-1",
trace_id: "trace-1",
metadata: { evaluator: "llm-judge" },
config_id: "cfg-abc",
data_type: "NUMERIC"
)
# Categorical string value
client.score(name: "label", value: "good", trace_id: "t1", data_type: "CATEGORICAL")# New v4 usage model (arbitrary keys, e.g. cache tokens)
gen = trace.generation(
name: "chat",
model: "gpt-4o",
usage_details: { input: 100, output: 50, cache_read: 30 },
cost_details: { input: 0.001, output: 0.003, total: 0.004 }
)
gen.end(output: "response")
# Link a generation to a prompt version (accepts a Langfuse::Prompt or a hash)
prompt = Langfuse.get_prompt("chat-prompt")
gen = trace.generation(name: "chat", model: "gpt-4o", prompt: prompt)
# or: prompt: { name: "chat-prompt", version: 3 }begin
client = Langfuse.new(
public_key: "invalid-key",
secret_key: "invalid-secret"
)
trace = client.trace(name: "test")
client.flush
rescue Langfuse::AuthenticationError => e
puts "Authentication failed: #{e.message}"
rescue Langfuse::RateLimitError => e
puts "Rate limit exceeded: #{e.message}"
rescue Langfuse::NetworkError => e
puts "Network error: #{e.message}"
rescue Langfuse::APIError => e
puts "API error: #{e.message}"
endclient = Langfuse.new(
public_key: "pk-lf-...",
secret_key: "sk-lf-...",
host: "https://us.cloud.langfuse.com", # US default; EU: cloud.langfuse.com
ingestion_mode: :otel, # Langfuse v4 (default remains :legacy)
debug: true, # Enable debug logging (or LANGFUSE_DEBUG=true)
timeout: 30, # Request timeout in seconds
retries: 3, # Number of retry attempts
flush_interval: 30, # Event flush interval in seconds (default: 5)
flush_at: 50, # Flush once this many events are queued (default: 15)
max_queue_size: 10_000, # Drop the oldest events beyond this many queued (default: 10_000)
auto_flush: true, # Enable automatic flushing (default: true)
environment: "prod", # Tracing environment (or LANGFUSE_TRACING_ENVIRONMENT)
sample_rate: 0.5, # Keep 50% of traces deterministically (or LANGFUSE_SAMPLE_RATE)
mask: ->(v) { v }, # Callable applied to input/output/metadata
shutdown_on_exit: true, # Flush pending events on process exit (default: true)
http_adapter: :net_http_persistent # Faraday adapter (default: Faraday's default)
)Timeouts, network errors, 429 and 5xx responses are retried up to retries
times (default 3). A Retry-After header is honored (seconds or HTTP date,
capped at 10 s); otherwise the delay grows exponentially from 0.5 s with ±50%
jitter, so many processes do not retry in lockstep. Client errors such as 401
and 422 are not retried, since repeating them cannot help.
By default every request opens a new TLS connection. To keep connections alive
between flushes, add a pooling adapter to your Gemfile and select it with
http_adapter:
# Gemfile
gem "faraday-net_http_persistent"
Langfuse.configure do |config|
config.http_adapter = :net_http_persistent
endIf the adapter is not available, the client logs a warning and falls back to Faraday's default adapter instead of raising.
You can also configure the client using environment variables:
export LANGFUSE_PUBLIC_KEY="pk-lf-..."
export LANGFUSE_SECRET_KEY="sk-lf-..."
export LANGFUSE_HOST="https://us.cloud.langfuse.com" # or LANGFUSE_BASE_URL; EU: https://cloud.langfuse.com
export LANGFUSE_FLUSH_INTERVAL=5
export LANGFUSE_FLUSH_AT=15
export LANGFUSE_MAX_QUEUE_SIZE=10000
export LANGFUSE_AUTO_FLUSH=true
export LANGFUSE_TRACING_ENVIRONMENT="production"
export LANGFUSE_SAMPLE_RATE=0.5
export LANGFUSE_DEBUG=false
export LANGFUSE_INGESTION_MODE=otel # v4; use `legacy` only for pre-v4 self-hostedBy default, the Langfuse client automatically flushes events to the server at regular intervals using a background thread. You can control this behavior:
# Enable automatic flushing (default)
client = Langfuse.new(
public_key: "pk-lf-...",
secret_key: "sk-lf-...",
auto_flush: true,
flush_interval: 5 # Flush every 5 seconds
)
# Disable automatic flushing for manual control
client = Langfuse.new(
public_key: "pk-lf-...",
secret_key: "sk-lf-...",
auto_flush: false
)
# Manual flush when auto_flush is disabled
client.flushLangfuse.configure do |config|
config.auto_flush = false # Disable auto flush globally
config.flush_interval = 10
endexport LANGFUSE_AUTO_FLUSH=falseAuto Flush Enabled (Default)
- Best for most applications
- Events are sent automatically
- No manual management required
Auto Flush Disabled
- Better performance for batch operations
- More control over when events are sent
- Requires manual flush calls
- Useful for high-frequency operations
# Example: Batch processing with manual flush
client = Langfuse.new(auto_flush: false)
# Process many items
1000.times do |i|
trace = client.trace(name: "batch-item-#{i}")
# ... process item
end
# Flush all events at once
client.flush# Ensure all events are flushed before shutdown
client.shutdown# config/initializers/langfuse.rb
Langfuse.configure do |config|
config.public_key = Rails.application.credentials.langfuse_public_key
config.secret_key = Rails.application.credentials.langfuse_secret_key
config.ingestion_mode = :otel
config.debug = Rails.env.development?
end
# In your controller or service
class ChatController < ApplicationController
def create
@client = Langfuse.new
trace = @client.trace(
name: "chat-request",
user_id: current_user.id,
session_id: session.id,
input: params[:message],
metadata: {
controller: self.class.name,
action: action_name,
ip: request.remote_ip
}
)
# Your LLM logic here
response = generate_response(params[:message])
render json: { response: response }
end
endclass LLMProcessingJob < ApplicationJob
def perform(user_id, message)
client = Langfuse.new
trace = client.trace(
name: "background-llm-processing",
user_id: user_id,
input: { message: message },
metadata: { job_class: self.class.name }
)
# Process with LLM
result = process_with_llm(message)
# Ensure events are flushed
client.flush
end
endCheck out the examples/ directory for more comprehensive examples:
- Langfuse v4 / OTEL tracing (recommended)
- Simplified usage
- Basic tracing
- Prompt management
- Events
- Auto-flush control
- Connection config
- Langfuse v4 usage — OTEL ingestion, observations-first model, cutover checklist
- Documentation index
- Publishing Guide
- Release Checklist
- Official Langfuse docs
- Migrate custom ingestion to v4
After checking out the repo, run:
bin/setupTo install dependencies. Then, run:
rake specTo run the tests. You can also run:
bin/consoleFor an interactive prompt that will allow you to experiment.
Bug reports and pull requests are welcome on GitHub at https://github.com/ai-firstly/langfuse-ruby.
The gem is available as open source under the terms of the MIT License.