Inspiration

We thought that building on the agent-to-agent ecosystem would be a very interesting project realm. Implementing x402, ANS, and Arweave seemed like perfect fits to contribute to the idea of transparency and decentralization within the agent-to-agent ecosystem.

This project also falls under the scope of "How do we scale AI without human bottlenecks?" which is also a very interesting realm to dive into.

What it does

Summary: Agent (with ANS identity) is given limits + demands for a custom model from user. Agent talks to a Vultr agent (system prompted model with context of Vultr documentation/API documentation, also with ANS identity) and negotiates and discusses ideal CPU plans for said model. Once both agents sign, the agent pays over x402 to rent the CPU instance, and uses it to train the model the user requested. The entire interaction, pass or fail, is uploaded to Arweave to ensure the record is untouchable, public, and permanent. A clean agent sits at a trust score of 100, and every anchored violation halves it: one failure drops it to 50, two to 25. We chose to halve rather than incrementally dropping because a single fail means it sucks at its job, and fails can lead to massive consequences when handling company (larger-scale) demands. The score is computed from the verdicts anchored on Arweave and attributed to the agent's ANS identity, so it can't be tampered with.

In depth: Your agent gets registered an ANS identity of its own, so that all records and transactions (and failures) can be attributed to it. You then give the agent training data, as well as what you require from a custom model, and set three limits: a total budget, an hourly spend rate, and a timeframe.

Your agent then opens an A2A conversation with a second agent that sells Vultr compute (named Vultr desk). Both sides sign every message and verify the other against keys tied to its ANS identity, sealed in a transparency log. The desk quotes real plans at real prices for CPU instances with timings for that exact job, and recommends plans. For example, in one of our runs our agent pushed for a cheaper four core box and the desk refused, pointing out that the job splits into sixteen parallel pieces and four cores would run them in four serial batches, AKA it would make the work take four times longer, which went against the original wishes of the user.

Once they agree, the agent pays the broker per request over x402 in USDC on Solana (Devnet for demo), and the broker provisions the instance. Only the broker holds the Vultr key, so no agent in the chain ever touches it. The box boots, trains, serves its own progress, and hands back a model as plain JSON that runs in your browser. Then the box is released, destroys itself, and the model keeps working without it.

Every transaction, approval, every refusal and every broken rule is signed and written to Arweave (Testnet or Mainnet). Anyone can re-run the audit from the public record with one command:

npm run burn402 -- verify <arweave-txid>

How we built it

Next.js for web platform, Arweave for storage, x402 for payment gateway, Ed25519 and JWS for signatures.

The mandate is the core: a signed permission slip with a budget, an hourly ceiling, a scope and an expiry, plus ten attenuation rules that the broker checks before anything is provisioned. The x402 gate answers 402, verifies the payment, provisions, settles and signs a receipt. A reaper watches the burn rate and destroys the instance once reaching zero.

The trainer on the box uses only the Python standard library: a hashed tf-idf classifier trained by gradient descent, a character n-gram language model, and a tf-idf passage index. All three start from scratch on your data. Nothing is pretrained (check it yourself!)

The auditor takes public evidence, runs its own eight checks over it, and signs a verdict that carries the evidence it was made from, so anyone can reproduce it. Verdicts feed the Trust Index as a behaviour score, and anchored records live on Arweave attributed to the agent's domain (ANS identity).

The auditor does not have to be told where to look. It lists every instance on the provider account and diffs that against the receipts that settled, so a server nobody paid for shows up on its own.

The negotiation itself is anchored too. Every message between the two agents is a signed JWS, and the exchange goes to Arweave alongside the payment receipt, so you can check afterwards what the desk actually recommended and whether the agent rented it. The desk's side is stored as text. Your side is stored as a SHA-256 commitment instead, so the request you typed and the name of your file never go public, and anyone holding the original message can still prove it is the one the record commits to. A toggle on the page publishes your side in full if you want it there.

Timings are measured. We rented one box per Vultr CPU family and timed the same job on each, then fitted a cost model to the shape of the data. The page shows predicted time against actual, so you can see whether the estimate was realized.

Web3 stuff

You don't request burn402 for a GPU/CPU instance, your agent actually buys one which is made possible with x402. The broker answers the agent's request with HTTP 402 and a price, the agent pays that price in USDC on Solana devnet through an x402 facilitator, and the machine only boots once the payment settles. Using Solana + Arweave is what allows us to create a receipt for the auditor to check later (along with Solscan).

The payments are SPL token transfers, and there is one per request instead of a single settlement at the end. A job costs between 0.0006 and 0.0165 USDC depending on which plan the two agents agree on, so a session is a handful of tiny machine-to-machine payments. Every one of them returns a transaction signature you can open on Solscan, and the broker signs a receipt containing that signature, the mandate it was paid against, and who paid. That receipt is what we anchor on Arweave, so the payment and the reason for it stay attached to each other.

The dashboard reads both wallets from the chain with getTokenAccountsByOwner, so you can watch the agent's USDC balance drop and the broker's climb while the job runs. Here is a real one from our GoDaddy documentation run: 0.011 USDC for a vhp-8c-16gb-amd instance, on devnet.

5MwMc15SMDKLZS88ACd83Fv4ACfdvZ1KNEYAR29B6pQMZd2QEtwMWbqaFYqbX7v1VDLdMSTS46bY2HA5vj485eXK

Challenges we ran into

Vultr has GPU instances blocked on our account until we file a support request to unlock them, so everything had to finish on CPU (this was a HUGE bottleneck). That ruled out fine-tuning anything pretrained and pushed us toward models that train from zero and still produce something that works. If we had GPU access we could have been much much much more ambitious with the kind of models we could have produced.

The delegation chain caused problems too. When we wired the stress test to attack a real mandate, our own attenuation rule rejected it for raising the maximum delegation depth above its parent.

Accomplishments that we're proud of

It runs on real infrastructure e2e. Real Vultr instances, real USDC on Solana devnet, real ANS identities stored in a transparency log, real records on Arweave mainnet. Nothing is hardcoded or faked for the demo (verify this yourself!)

Verdicts are reproducible. The record carries its own evidence and a hash of it, so you do not have to just take our word for it:

npm run burn402 -- verify <arweave-txid>

downloads it, checks the auditor's signature against the logged key, and runs every check.

You can audit the reasoning as well as the outcome. The signed exchange between the two agents goes on Arweave with the receipt, so a judge can check that the plan we rented is one the desk actually named, and that the agent did not quietly pick something else. verify accepts those records and checks both agents' signatures against the transparency log:

The stress test attacks the budget YOU actually set. It reads your hourly cap, finds the cheapest plan above it, asks for that, and gets refused with your own numbers thrown back at it. Then it cheats: it skips the broker entirely and rents that box straight from Vultr, and nothing stops it, because nothing can stop an agent that already holds a provider key. What happens next is the part we care about. The auditor lists the provider account itself, diffs it against the receipts that settled, spots a server that nobody paid for, and signs a verdict that goes on Arweave permanently and halves the agent's trust score. The agent got its server. It also got a permanent record saying it broke the rules, attached to its identity, which anyone can look up.

What we learned

First usage of x402, Arweave, and ANS. This was our first project involving agent-to-agent commerce but we always thought it was a super interesting field so learning about how it works was really cool.

We also learned to measure instead of guess. Our first cost model was made up, and the CPU families turned out to differ far more than we assumed: the same training job took 19.6 seconds on one high frequency core and 54.6 seconds on a regular one. That is why the timings are now measured per family, and why the page shows predicted against actual instead of asking you to trust the estimate.

What's next for burn402

GPU jobs when Vultr unblocks our account. We could try fine-tuning open weight models like Qwen 2.5 on a rented GPU but it would also require changes in scope to the mandate and a different trainer.

Built With

  • a2a
  • agent-to-agent
  • ans
  • arweave
  • solana
  • x402
Share this project:

Updates

Submission history