We will be undergoing planned maintenance on Oct 7th 6:00AM UTC / Oct 7th 2:00AM ET

Who it's for: a parent buying used baby gear from a stranger online and paying for it at the curb. What they get: the money doesn't move until a camera has read the item's label and checked it against CPSC nursery and children's recalls, NHTSA car seat recalls and the federal resale bans. If the item is recalled or banned, the Visa hold gets reversed and the parent never paid.

Inspiration

I spent over four years working with kids ages one to five in child care and camp programs. A lot of parents would show up with a secondhand stroller or a bassinet they got off Facebook Marketplace, and there was really no easy way for them to know if it had been recalled. Most of them had no idea recalls even applied to used stuff.

And it's not a small problem. About 100 babies died in Fisher-Price Rock 'n Play sleepers, and at least 8 of those deaths happened after the recall (CPSC, Jan 2023). This year Consumer Reports looked at Facebook Marketplace and found 50 of the first 65 listings it reviewed were banned infant in-bed sleepers, before it stopped counting.

Marketplaces filter what they can read in a listing. Nobody reads the label on the actual item when the money changes hands in a parking lot. So that's the moment we built for.

What it does

The whole idea fits in one line: the money doesn't move until the camera has seen the item.

Find it. The parent types or just talks to our shopping agent ("a bassinet for my newborn under $80, pickup in Atlanta"). Gemini turns that into a search over 1662 real listings we scanned, and every result comes back already checked. Red ones can't be bought, amber ones need a look at pickup. You can also do the whole thing by voice with our ElevenLabs agent. Agree and hold. The buyer agrees on a price and types their card into Visa's own Microform fields. Lullabuy asks Visa to authorize it with capture off, so the money is held, not sent. Meet and scan. At pickup the buyer photographs the label. The barcode gets read on the phone, the label text gets read into model, batch and manufacture date, and our own model looks at the photo for product types that are banned outright (inclined sleepers, padded crib bumpers, drop-side cribs). Check. Those values go against 1136 CPSC nursery and children's recalls and 71 NHTSA child car seat recall campaigns. The rules follow the law: a recall only counts for the batch or date range it names, mesh crib liners are allowed, and a car seat always gets NHTSA's used-seat check. Capture or reverse. No match, the hold is captured and the seller gets paid. Recalled or banned, the hold is reversed. Anything uncertain keeps the hold and moves no money. The app never says "safe". It says "no recall match as of" the date it checked. Get the call. This is the part people react to. If the hold gets reversed, Lullabuy actually calls the parent. A real voice (ElevenLabs) tells them which CPSC recall their item matched and that they weren't charged, and they can press 1 to hear it again. It only ever calls a phone that said yes: the first call just reads a 4-digit code that you type back in. Or settle it at a table. At a swap meet or a consignment sale, the seller plugs in a regular USB barcode scanner and opens /checkpoint. One scan of the UPC settles the hold. If the scan comes in broken, or a recall's UPC is off by a digit, it doesn't guess. Nothing moves until a person checks. After the sale. Every captured sale gets an item passport on Solana that anyone can verify, and it's also minted as a Metaplex Core asset. The recall watch keeps checking every sale. If a recall comes out later for something a parent already bought, they get a push notification and the asset's on-chain attributes get updated.

Who it helps, and why a marketplace would pay for it

This is a big market. GlobalData, in a report for Mercari, valued the U.S. secondhand kids and baby market at $7 billion in 2021 and projected $12.8 billion by 2030 (that covers kids' items broadly, clothing included, not just cribs and car seats).

The parent never pays for a recalled or banned item, and they don't have to know the recall list. One photo at pickup does it.

The marketplace gets fewer disputes, because a reversed hold means there's no charge to dispute. Visa's own 2026 report puts a merchant's average cost to resolve a first-party misuse dispute at $82 (Visa/MRC 2026). It also gets less legal exposure: selling a recalled product is against federal law, and the maximum civil penalty is $120,000 per knowing violation (Federal Register, 2021).

So the buyer for this is a marketplace's Trust and Safety and payments team, and /trust is the screen they'd actually watch.

How we built it

Visa Acceptance (sandbox): authorization with capture off, full capture and full reversal, signed with HTTP Signature, running on our own Visa sandbox merchant. A signed deal token binds the deal, the authorization and the amount, so the pickup scan can only settle the hold it was issued for. Replaying a settlement is refused by Visa and shown as refused. Visa Acceptance Microform: the card number and CVV are typed into Visa-hosted fields and come back as a one-time token, so the card never touches our server. Verified end to end in a browser against the sandbox. Visa Token Management Service and card-linked offers: a buyer can save the card with Visa and pay the next hold without typing it again. The browser only keeps our signed wrapper around Visa's token. While testing it we found that when Visa applies a card-linked offer ("20 percent off" a card's 10th purchase), the reply carries a lower authorized amount and no status field. We hold exactly what Visa authorized and show the saving. Trusted Agent Protocol: when our AI agent buys for a parent, it signs its checkout with RFC 9421 HTTP Message Signatures (Ed25519), covering the method, host, path and a digest of the body. Our merchant checks the signature, the time window and a one-time nonce before calling Visa. On /pickup you can send a request whose amount was edited after signing and watch it get refused. Verified means the agent is who it says and the request wasn't changed. It doesn't mean a parent approved the purchase. ElevenLabs: a voice shopping agent built on ElevenLabs Agents that searches our real listings through a webhook tool, plus the verdict spoken out loud at the curb in English or Spanish. Recall phone call (Vonage Voice + ElevenLabs): when a hold gets reversed, or the recall watch flags a sale later, Lullabuy places one outbound call and plays the verdict in an ElevenLabs voice. If ElevenLabs fails, the call still happens with Vonage's own voice. The number gets checked first with the code call, it's sealed with AES-256-GCM, and it never shows up on a public page (the buyer's own page shows the last 4 digits). There's one call per deal, a cap per number and a daily cap, all enforced with atomic writes in Atlas. We tested it on our own phones tonight, start to finish on the live site. Table kiosk (/checkpoint): a USB barcode scanner types like a keyboard, so the kiosk just reads what it types. Every code has to pass its GTIN check digit, a paused or partial scan gets thrown out, and if two devices scan the same deal at once, an atomic claim in Atlas makes sure Visa only settles it once. Gemini: Gemini 3.5 Flash turns a spoken or typed request into filters over our scanned listings, and it reads the label photo into brand, model, batch and date with a box for each field. It never decides what's allowed. The recall index, our model and our own review do, and the server refuses to let the agent buy a red listing. In production Gemini runs with no stored key: Vercel's identity token is exchanged with Google Workload Identity Federation for a short-lived token. Recall index: 6036 CPSC recalls pulled from the CPSC API, 1136 kept as nursery and children's products, plus 71 NHTSA child restraint campaigns. Gemini read all 1136 recall notices for model numbers, batches and UPCs, and a value is kept only if it shows up word for word in that recall's text. That raised the recalls we can match on from 449 to 820. Two hand checks of 20 records each found three extraction bugs, which we fixed and wrote up in data/handcheck.md. Banned-type model: CLIP ViT-B/32 image embeddings with a logistic-regression head we trained on reviewed labels. It runs right in the buyer's browser with transformers.js, and we trained it on embeddings from that same runtime so the phone sees exactly what training saw. On products it had never seen, macro-F1 is 0.7243 (95% CI 0.6017 to 0.8127), against 0.558 for off-the-shelf CLIP. On ordinary baby items it wrongly flags 2 of 107, against 38 for off-the-shelf CLIP. The Atlanta scan: we harvested 1662 real listings (Craigslist Atlanta and eBay), ran every one through the model, then read every flag ourselves. Round 1 raised 39 flags. We added the false alarms (crib skirts, rail covers, dollhouse cribs) to training, without touching the test set, and round 2 raised 9. MongoDB Atlas: every hold and settlement is recorded. A change stream pushes each update to the deal board the moment it happens, the seller scans a QR on the buyer's phone and watches the same deal live, and Atlas Vector Search matches the photo embedding against CPSC recall photos. Tested on held-out photos (a recall's second photo, searched against an index holding only its first), the right recall came back first 60 of 144 times and in the top three 68 of 144 times. A look-alike is only ever a reason to read the label. Trust and Safety console (/trust): the view a marketplace's safety and payments teams would actually watch, computed live from Atlas: money that never reached a seller of a recalled item, reversals grouped by the actual CPSC recall number, and time from hold to decision. Recall watch: every captured sale keeps the label its scan read, a daily job re-checks all of them against the recall index, and the owner gets a web push if something they already bought gets recalled. On /watch a judge can type a model number and see which real sales would get caught. Solana item passport: when a sale captures, a Memo transaction on Solana devnet stores the SHA-256 of the pickup record (no personal data on chain), and the sale is minted as a Metaplex Core asset with the check result as on-chain attributes. The recall watch updates those attributes if the item gets recalled later. The passport page reads it back and checks the signer, the transaction and the hash, so anyone can verify what the check said at the time of sale. Hold sweeper: a daily job releases any hold whose pickup never happened, so a parent's money isn't left held for more than about a day. MCP server: the recall check is also an MCP tool (recall_check at /api/mcp), so any AI shopping agent can check an item before it buys. Mobile app: the same product on iOS and Android (Expo): shop, scan a label with the camera and the deal board, all on the live API. The Android APK is on the GitHub release mobile-v1.0.0. Frontend: Next.js 16, GSAP and Lenis motion, MapLibre with OpenFreeMap for the scan map. Engineering: CI runs lint, types, unit tests, a build, pytest, a JS-vs-Python classifier parity suite, Playwright end to end, a secret scan and an em dash check. A probe checks both production URLs on a schedule.

How we used Notability

We planned and rehearsed the whole build in Notability. Our architecture, money flow and expo demo note lives there, and Notability's AI turned it into Smart Notes ("System Architecture & Payment Flow") and a 20-question practice quiz that we actually used to drill judge questions, like "what triggers a reversal of the payment hold?". Screenshots are in the gallery and in docs/stills/notability/.

Challenges we ran into

The model's first pass on real listings flagged crib skirts and dollhouse furniture. Reading every flag and feeding the mistakes back fixed most of it. Banned products are rare on eBay and Craigslist (eBay filters them out), so our test set had to be grouped by product instead of being made only of listing photos. Our first recall extraction turned phrases like "4-in-1" and years into model numbers. The listing scan exposed it. CPSC's own API sometimes pairs a recall with the title of a different recall. One record showed a Joolz stroller adapter title next to a CooCooBaby lounger notice, so a Joolz model number matched the wrong recall. We now take the title from the recall page itself whenever the two disagree, and a test checks every record. Payment states: an adversarial review found that a failed or replayed settlement could show up as HELD. Now anything Visa didn't confirm is shown as refused or unknown, never as a result. Gemini's recall pass added model numbers like "4340" that several brands use. Our first fix trusted the brand if the listing named it, and review found brands like "Summer" and "Gap" in ordinary listing text. Now a match on a short number never moves money: the hold waits for a person to check the brand. Barcode scanners type really fast, and our first kiosk tried to be smart about stray keys. Review showed the end of a recalled code can be a totally different clean barcode (04871507 sits inside 698904871507), so a paused scan could have paid for a recalled item. Now Enter sends the whole scan or nothing.

Accomplishments that we're proud of

A real Visa hold that a camera settles, end to end, tested in a browser on the live site. A reversed hold that picks up the phone and tells the parent, out loud, which recall it was. A model that runs on the buyer's phone and was measured on products it had never seen. Every number on the site is recomputed by a script and served from an open endpoint, so you don't have to take our word for any of it.

What we learned

The danger isn't where we first looked. We expected banned sleepers all over eBay. Our scan of 1662 real listings found about one. The big marketplaces already filter the text they can read, so the real problem is at the handoff, when cash or a card moves for an item nobody has looked at closely. A model is only as good as the mistakes you feed back. Our first pass flagged crib skirts and dollhouse furniture. We reviewed every flag by hand, added the false alarms to training, and the flags went from 39 to 9 without touching the test set. Money code needs someone actively trying to break it. Review rounds found bugs our tests didn't, like model numbers such as "4340" that several brands share, and brand names like "Summer" and "Gap" that show up in ordinary listing text. We changed the rule so a short number never moves money on its own. Visa's hold is honestly the right tool for this. Authorize with capture off, then capture or reverse, is how hotels and gas pumps already work, and it fits a parking lot sale exactly. You can use a cloud model without a stored key. Our server proves who it is to Google with a short-lived identity token, so there's no Gemini key sitting in our settings.

What's next

Seller-side confirmation of the scan (today the buyer's device reports it). Paying the seller right away with Visa Direct once the hold captures. The code is written and tested, but it isn't turned on in this deployment yet because it needs Visa Direct access we don't have tonight. Bring the check to Facebook Marketplace handoffs, where the problem is the worst.

Every parent I worked with deserved to know if the thing they were buying for their kid had been recalled. Now the money just waits until somebody checks.

Try it

Live: https://lullabuy.tech (judges, start at /judge). Also at https://secondhand-safe-web.vercel.app. curl "https://lullabuy.tech/api/check?model=BHC001&batch=202408" MCP: {"mcpServers":{"lullabuy":{"type":"http","url":"https://lullabuy.tech/api/mcp"}}}

Built With

  • atlas-vector-search
  • clip
  • cpsc-api
  • elevenlabs
  • expo.io
  • gemini
  • github-actions
  • maplibre
  • mcp
  • mongodb-atlas
  • nextjs
  • notability
  • playwright
  • python
  • react-native
  • scikit-learn
  • solana
  • transformers.js
  • trusted-agent-protocol
  • typescript
  • vercel
  • vertex-ai
  • visa-acceptance
  • visa-microform
  • visa-token-management-service
Share this project:

Updates

Submission history