Palmyra X6

PALMYRA X6

Built for the long run.

Our flagship model for agentic work at scale. It executes complex work for hours across the customer journey — and it’s efficient enough to run from one marketing team to your whole revenue org.

X6
Vanguard
Salesforce
KPMG
Uber
e.l.f.
SCAN Health Plan
VOIS
Vanguard
Salesforce
KPMG
Uber
e.l.f.
SCAN Health Plan
VOIS
Trophy

Benchmarks

See how our model stacks up

1

0.87

Top capability score
in a six-model field

$

$0.12

Average cost per finished task.
52% less than the last generation.

Time

8 hrs

On a single objective,
start to finish, unattended

WHAT WE MEASURE

We didn’t train it on trivia.
We trained it on the work.

Every score here comes from the production work our customers actually run — marketing, revenue, research, outreach — graded 0 to 1. Palmyra X6 improved on all nine capabilities we measure.

A performance scorecard on a solid dark charcoal background, presented as a nine-row horizontal metrics table with four unlined columns. Each row represents a WRITER AI capability, shown from top to bottom in this order:

Sub-agents — Delegating work to spawned sub-agents and merging results — score 0.86 — change +0.12
Grounding & retrieval — Knowledge-graph and document grounding, retrieval, and citation — score 0.91 — change +0.10
MCP tool use — Discovering and calling external tools through MCP connectors — score 0.88 — change +0.10
Content generation — Drafting and editing long- and short-form content to a brief — score 0.86 — change +0.08
Playbooks — Executing predefined multi-step workflows end to end — score 0.84 — change +0.08
Model & system awareness — Knowing its identity, tools, scope, and when to safely refuse — score 0.97 — change +0.02
Brand voice — Applying brand voice and custom instructions to the output — score 0.92 — change +0.04
Image analysis & gen — Interpreting input images and generating new images — score 0.80 — change +0.09
Presentations — Generating slide presentations from a prompt or source — score 0.74 — change +0.10
The table has four distinct columns: the far-left column contains short category titles in white, medium-weight sans-serif font; the middle-left column contains longer descriptive text in a smaller, lighter gray sans-serif font; the middle-right column contains horizontal progress bars with a black track and a bright cyan/light-blue fill whose length visually represents the score; and the far-right column contains two numbers — the base score in white and the positive delta (marked with a "+") in the same bright cyan as the bar fill. Row spacing is uniform, creating a clean, scannable list. The bright cyan bars and delta numbers provide the strongest visual contrast against the dark background, drawing immediate attention to the performance metrics and their improvements.

ENDURANCE

Eight hours. One goal.
No hand-holding.

Most models lose the thread in minutes. Palmyra X6 holds one goal for up to eight hours — and hands back finished work, not a draft.

A horizontal timeline diagram on a solid dark charcoal background illustrating the chronological progression of an automated WRITER workflow over an 8-hour period. At the top, four side-by-side phase blocks are labeled, from left to right: PLAN, EXECUTE, OPTIMIZE, and DELIVER. The blocks use a monochromatic color progression — PLAN is dark indigo/purple, EXECUTE is a vibrant medium blue, OPTIMIZE is a lighter cornflower blue, and DELIVER is a very pale lavender/white. EXECUTE is the widest block, spanning from 01:00 to 05:00.

Directly below the phase blocks is a horizontal timestamp axis with evenly spaced labels reading: 00:00, 01:00, 02:00, 03:00, 04:00, 05:00, 06:00, 07:00, 08:00.

The main body contains five white, rounded-rectangle event cards arranged in a cascading diagonal pattern from top-left to bottom-right. Each card contains a timestamp and an event description:

00:00 — goal received
01:00 — plan locked • 41 tasks
03:42 — self-test failed -> fixed
06:15 — throughput +34%
08:00 — shipped • verified
The text inside the white cards is black, except for the final "08:00" timestamp, which is colored in the same pale lavender as the DELIVER block. Faint thin vertical lines drop down from specific timestamps on the axis (00:00, 04:00, 06:00, 08:00) to align near the white cards. At the bottom left, the text "8 hours" sits above a thin white horizontal line that spans the width of the image, ending in an arrowhead pointing to the right. The timestamps and card details use a technical monospace typeface, while the phase headers and "8 hours" text use a clean sans-serif typeface.

HOW IT COMPARES

Top-performing,
at a fraction of the cost

We plotted every model by what it does and what it costs. Palmyra X6 scores highest — at far less than the other frontier models charge.

Scatter plot chart on a dark grey background comparing AI models on two metrics: Capability Score (Y-axis) and Cost (X-axis).

The Y-axis on the left is labeled "CAPABILITY SCORE" with values "0.65," "0.70," "0.75," "0.80," "0.85," "0.90." The X-axis along the bottom is labeled "COST PER 1M OUTPUT TOKENS" with values "$1," "$5," "$10," "$20," "$30."

Data points are scattered across the plot, represented by white circles, except "Palmyra X6," which is represented by a distinct larger stylized-face icon. Labeled models include "Palmyra X6," "Claude Sonnet 4.6," "GPT-5.5," "Gemini 3.1," "Claude Opus 4.8," and "Flash 3.5."

A highlighted rectangular zone in the top-left quadrant (high capability, low cost) is marked with a vertical "BEST VALUE" label. This highlight zone, the "BEST VALUE" text, and the Palmyra X6 icon all use a distinct mint green/teal color to draw attention.

9.4x

Cheaper output than
Claude Opus 4.8

Top
score

highest capability at the lowest frontier price

$3.50

Blended / 1M at a typical
3:1 input-to-output mix

BIAS & NEUTRALITY

Even-handed,
where it counts most

We asked every model the same hot-button political questions and scored whether it argued one side or both. Palmyra X6 came back the most even-handed — well ahead of every other frontier model tested.

A horizontal bar chart comparing AI models on the percentage of answers that present both sides of an issue. The chart is set against a dark blue and purple textured background resembling woven fabric, and the chart itself sits inside a white rounded rectangle.

The model Palmyra X6 is the focal point at the top, with its name in larger, bolder text and a bright cyan/aqua bar reaching 80%, with the percentage label in dark teal. All other models below it are shown in standard black text with muted grayish-purple bars and medium-gray percentage labels.

The models and their values, from top to bottom, are:

Palmyra X6: 80%
Claude Opus 4.8: 57%
GLM-5.2: 43%
Kimi K3: 39%
Mistral Medium 3.5: 34%
GPT-5.6-sol: 33%
DeepSeek V4: 17%
Gemini 3.6 Flash: 9%
The X-axis along the bottom runs from 0 to 100 with gridlines at 20, 40, 60, 80, and 100, and is labeled "% OF ANSWERS PRESENTING BOTH SIDES."
Multi-model support

MULTI-MODEL SUPPORT

Flexible model support
for when you need it.

Palmyra X6 leads our evaluation — but you’re not locked in. Run the model each job needs, from any provider, or bring your own.

Marketing UI mockup illustrating an administrative settings panel titled "Set what's available," with the subheading "Control what models your teams can access." The image is set against a dark, grainy gradient background of black and deep purple.

Three overlapping white UI cards are arranged in the center. The foreground card on the right is labeled "WRITER" with an "8 of 8 models enabled" status and a "Disable all" link in purple text. A table-style layout shows column headers "MODEL NAME" and "ENABLE." Listed models include "Palmyra X6," which has a light purple pill-shaped badge reading "Default agent" with a user profile icon, and "Palmyra X5." Two purple toggle switches are shown in the active/enabled state. An upward-pointing chevron indicates a collapsible menu.

Behind it on the left is an "Anthropic" card showing "9 of 9 models enabled" with column headers "MODEL NAME" and listed models "Claude Sonnet 4.6" and "Claude 3.7 Sonnet," with toggles in the active state, slightly faded. A bottom card partially visible on the lower right is labeled "Google" with "0 of 5 models enabled." Each provider card displays a corresponding logo icon.
Marketing UI mockup demonstrating a chat interface model selector, titled "Pick per session," with the subheading "Users can pick the model they'd like to use for each session." The background is a dark, textured purple-black gradient.

Below the heading is a horizontal white chat input box containing the placeholder text "Ask anything." The input box includes interactive icons: a plus sign for attachments, a model selector button showing "Palmyra X6" with an up chevron indicating the menu is open, a prompt library icon, a microphone icon, and a black circular send button with a white arrow.

Emanating from the model selector is a white dropdown menu on the left with a search bar at the top containing a magnifying glass icon. The menu lists "Palmyra X6" with a "Default" badge and checkmark indicating the currently selected model, "Palmyra X5," "Claude Sonnet 4.6" with a light purple background indicating a hover/focused state, and "Claude Sonnet 3..." cut off.

On the right, overlapping the dropdown, is a detail card for the focused model "Claude Sonnet 4.6" by Anthropic. It reads: "Excels at complex codebase navigation, end-to-end project management with memory, polished document creation." Below is the pricing "$3/1M Input tokens."
Marketing UI mockup showing a configuration screen for adding third-party custom models, titled "Bring your own," with the subheading "Bring custom models through cloud providers." The background is a dark, grainy gradient of black, purple, and blue.

A single white UI card is centered below the heading. The card header reads "Add provider." Inside, a horizontal two-step progress stepper displays "SELECT PROVIDER" as the active state (dark text, purple circle with a minus symbol) and "CONFIGURE MODEL" as the pending state (light grey text, empty grey circle), connected by a thin line.

Below the stepper is a vertical list of three selectable provider buttons, each with a corresponding logo icon on the left: "AWS Bedrock" in an active/selected state with a purple border, "NVIDIA NIM" and "Baseten" in an unselected state with light grey borders.
Marketing UI mockup showing a settings panel for managing specialized AI models, titled "Specialized models," with the subheading "Preset models for specialized tasks like image generation." The background is a dark, grainy gradient of black and purple.

A single white UI card is centered below. The card header reads "Image generation models" with "2 models enabled" and a pill-shaped "+ Add model" button with a white background and grey outline aligned to the right.

Below the header is a two-column table with headers "MODEL NAME" and "DEFAULT." Two rows are listed:

"GPT-5" with a selected purple radio button (purple circle with white dot) and a light pink pill-shaped badge reading "Default image" with a user icon.

"Gemini 3 Pro Image Preview" with an unselected empty grey radio button.

Each row features a drag-handle icon (six dots) on the far left, a provider logo, the model name, a radio button, and a trash can delete icon on the far right. The unselected row's text is a slightly lighter grey.

ENTERPRISE-GRADE

Complete IT control,
without the governance overhead.

Your data never trains our models. Guardrails can screen for PII and copyright before anything ships. And Palmyra X6 works inside the tools and policies your teams already run — so governance travels with the work, everywhere it goes.

→ Explore AI Studio

Padlock

Your data stays yours

Prompts, completions, and documents never train our models. Zero retention by default. Self-hosting and full data residency are available.

Guardrails

Guardrails

Configurable guardrails detect and redact PII and screen outputs against copyrighted material — before a single word leaves the building.

Governed & controlled

Governed & controlled

SSO/SCIM, granular RBAC, audit logs, customer-managed encryption keys, and alignment to the NIST AI RMF and the EU AI Act.

Frequently Asked Questions

FAQs about models

What model powers WRITER?

Palmyra X6 — WRITER’s most capable model — is the default model across the WRITER platform. It’s built for the work marketing and revenue teams actually run: grounded in your company knowledge, able to hold your voice, able to call the tools your stack already runs on, and coherent across hours of work rather than minutes.

X6 shipped alongside a new version of WRITER Agent. The two were developed together — the model was built to run inside the agent, and the agent was built around the model.

Is Palmyra X6 the only model I can use?

No. X6 leads our evaluation, so it’s the default — but it isn’t the only option.

Admins can enable models from other providers, and users can choose a model at the start of a session. Specialized models can be selected for specific capabilities, such as image generation. And beyond WRITER’s own catalog, admins can bring their own models from AWS Bedrock, Microsoft Azure, and NVIDIA NIM.

How was Palmyra X6 evaluated?

Public benchmarks measure what’s easy to grade. They don’t measure whether a model can pull twenty-two sources through a knowledge graph and cite each one, hold a voice across a 3,000-word draft, or hand half a research task to a sub-agent and merge the results back correctly.

So we built the evaluation out of the work itself — production tasks our customers run across marketing, revenue, research, and outreach — graded 0 to 1 against a locked baseline across nine dimensions, each chosen because it’s a common failure point in enterprise deployments. X6 improved on all nine dimensions, with no regressions.

The scores measure Palmyra X6 running inside WRITER Agent — the model plus the system around it — because that’s how customers run it.

How does X6 compare to other frontier models?

We ran the same evaluation against five other frontier models. X6 recorded the highest capability score in the set at the lowest frontier price. Models priced two to nine times higher scored no better on this work, and the one materially cheaper model scored considerably worse.

What does it cost to run?

X6 costs $2 in / $8 out per million tokens. It finishes the average task for roughly $0.12 — about half the per-task cost of the previous generation, in about half the time. Output tokens cost 9.4x less than the most expensive frontier model in our comparison set.

Per-task cost is the number that matters, because the unit of work has changed. It used to be a prompt and a response; now it’s an objective handed to an agent that plans, researches across twenty sources, drafts, checks its own work, and returns something finished. A model priced for occasional high-stakes questions becomes a different proposition entirely for a marketing team running four hundred campaigns a quarter.

How long can an agent run on a single objective?

Running inside WRITER Agent, X6 holds one goal for up to eight hours — planning, executing, testing its own output, correcting, and delivering finished work rather than a draft someone has to supervise.

Most models drift off a long objective within minutes. The difficult part is holding intent, not just context, across hours rather than turns — and it’s what makes background automation practical: agents left running to watch for market signals or competitor moves and act when something happens. Because each step is efficient, those runs are inexpensive enough to leave on.

Do you train on our data?

No. Your data never trains our models. Prompts, completions, and documents stay yours, with zero retention by default. The playbooks, workflows, and institutional knowledge your teams build remain yours.

Guardrails screen output for PII and copyrighted material before it leaves the platform. Administration includes SSO/SCIM, granular RBAC, audit logs, and customer-managed encryption keys, aligned to the NIST AI RMF and the EU AI Act. WRITER holds SOC 2 Type II, ISO 27001, ISO 42001, HIPAA, GDPR, and PCI-DSS. Self-hosting and full data residency are available.

How do you handle political and ideological slant?

Marketing and revenue content goes out under your brand, to customers and prospects — so slant isn’t hypothetical, it’s a brand risk, and one that’s hard to catch by hand once you’re generating thousands of pieces a month.

We evaluated Palmyra X6 against seven other frontier models across eleven benchmarks, including the Washington Post’s ModelSlant eval. X6 presented both sides of a hot-button political question in 80% of its answers — the most even-handed result tested, well ahead of the next-best score of 57% — while also posting the field’s second-lowest refusal-mismatch rate at 1.2%. Full methodology and per-model results are in the transparency report.