Inspiration
We were inspired by the number of vulnerable people being hurt by advances in scamming. In 2025, Americans aged 60+ filed 201,266 complaints with the FBI and reported $7.7 billion in losses, a 59% jump in a single year. They made up 20% of complaints but 37% of all money lost, with an average loss of about $38,500, and more than 12,400 seniors lost over $100,000 each. Scams are also getting smarter: seniors filed over 3,100 FBI complaints referencing AI in 2025, including voice-cloned "grandchild in distress" calls.
Phone calls are where older adults get hurt the most per incident. According to the FTC, scams that start with a phone call carry the highest median loss for older adults: $2,210, more than three times the $650 median for scams that start on social media. In 2024, 41% of older adults who lost $10,000 or more to a business or government impostor said it started with a phone call. And these reported numbers are only the tip of the iceberg: the FTC estimates the true cost of fraud to older adults in 2024 was between $10.1 billion and $81.5 billion.
The script is almost always the same: someone official-sounding, a sudden crisis, a demand for gift cards or wire transfers, and "don't tell anyone." These scams work because the victim is alone on the call. For the phone calls that make it through, Guardian Loop gives families the power to step into the conversation while it's happening, not after the money is gone.
What it does
Guardian Loop is a second pair of ears on a loved one's phone calls. It listens to a live call, transcribes both sides in real time, and scores the conversation from 0 to 100 across 10 scam signals: 7 from the caller (impersonation, threats, urgency, secrecy, untraceable payment, remote access, credential requests) and 3 from the victim (complying, disclosing personal info, or pushing back).
When the risk score hits 70, a trusted guardian (an adult child or caregiver) gets a notification short enough (under 180 characters) to read in full on a lock screen:
⚠️ Possible scam call (risk 82) "Buy the gift cards and don't tell your daughter." Why: untraceable payment request + secrecy request.
Tapping it opens a live dashboard with the transcript streaming in, red-flag phrases highlighted, the risk score charted over time, and a plain-English explanation, so the guardian can intervene during the call. Every call is saved to a History tab for replay.
How we built it
Guardian Loop is a ~9,500-line TypeScript monorepo running as a single Node service, with modules communicating over a typed in-process event bus.
- Calls: Two browsers connect peer-to-peer over WebRTC on an iPhone-style call screen. Each phone also streams its own microphone to the server in 20 ms frames, so we always know who's speaking.
- Transcription: Each speaker gets a dedicated Deepgram Nova-3 streaming session, with partial transcripts arriving in under a second.
- Hybrid detection: A fuzzy, speaker-aware rules engine (~40 rules using Damerau-Levenshtein matching, so "gift cart" and "Medicaire" still match) flags keywords instantly. Gemini reads a rolling window of the last 60 seconds (up to 10 turns, plus a one-line memory of the call so far) to catch scams that never use a keyword and to recognize innocent contexts, like a grandson mentioning a birthday gift card.
- Score engine: A pure, deterministic reducer combines both signals. 9 textbook combos (like impersonation + gift cards) add bonuses, and 7 of them set a sticky floor at or above the alert threshold. Each LLM read moves the score 50% toward Gemini's estimate, and the score decays 1 point per second when nothing suspicious is happening.
- Dashboard: React + Vite with no UI libraries, fed live over WebSockets, with browser notifications for alerts.
We wrote shared type contracts first so five workstreams (audio, core, AI, frontend, plumbing) could build in parallel against mocks, and kept 238 tests across 13 files passing in CI on every PR.
Challenges we ran into
- Two people, one room. During the demo both speakers sit side by side, so each mic hears both voices. We fixed speaker attribution in three layers: wired earbuds (about 16–20 dB of separation from mouth-to-mic distance alone), deliberately disabling automatic gain control, and a server-side gating automixer modeled on professional AV mixers. Every 20 ms it mutes the quieter track when the other is louder by 9 dB or more, with a 300 ms hangover so words don't get clipped.
- Speech-to-text errors. Exact keyword matching missed transcription mistakes, so we built order-independent, edit-distance-tolerant matching. The error budget scales with word length (0 edits for words of 3 letters or fewer, 1 up to 6 letters, 2 beyond) so short words don't trigger false positives.
- Phrases split across segments. "Read me your…" and "Social Security number" sometimes arrived as separate lines, so we pass the speaker's previous line (if under 5 seconds old) into the rules.
- The scammer controls the AI's input. Our LLM reads words spoken by the attacker, so we hardened it against prompt injection: the transcript is escaped JSON inside delimited blocks, and anything addressing "the AI" or "the monitor" is treated as evidence of a scam. The LLM can only set a score floor after two consecutive confident reads (80+), so one manipulated response can't trigger an alert.
- Keeping it real-time. Every Gemini call has a 3-second hard timeout and fails soft, triggers are debounced by ~1.2 seconds, and only one request runs per call at a time, so a slow API never stalls an alert.
- Browser and API constraints. Phone mic access requires HTTPS (solved with a Cloudflare tunnel), and a Gemini model retirement plus an API parameter change mid-hackathon forced us to adapt our classifier on the fly.
Accomplishments that we're proud of
- A full end-to-end pipeline from live call audio to a guardian alert in seconds.
- Proof the AI layer matters. We ran the same 87.5-second Medicare gift-card scam through the pipeline twice, once with rules alone and once with Gemini. With Gemini:
- Risk reached "elevated" at 21.9 seconds; with rules alone it never did, jumping straight from low to high at 49.8 seconds. That's 28 seconds of early warning.
- The score 22 seconds in was 52 instead of 19, because Gemini recognized "your Medicare account has been suspended" as a threat.
- It identified 8 of 10 scam signals instead of 3.
- It caught lines with no keywords at all, like "If you hang up, the case goes to the federal courts" and "Let me get my purse and my car keys."
- Average LLM response time was 1.09 seconds (range 866–1,226 ms), under our 1.5-second target.
- Explainable scores. Every number comes with a readable reason. Rules alone produced "impersonation + untraceable payment request"; with Gemini, the guardian reads "Caller impersonating Medicare demands gift cards to reinstate suspended benefits."
- Hybrid detection that can't be talked down. Rules set a floor the LLM can't erase, while the LLM catches what keywords miss.
- Graceful degradation. The system keeps running without Gemini, without a connected dashboard, or through a Deepgram outage (buffering about 10 seconds of audio while it reconnects).
- 238 automated tests and a pipeline testable without any vendor API keys.
What we learned
Real-time systems need to fail soft at every step; one slow API call can't be allowed to stall a live alert. Our hardest audio problem was solved with hardware first and code second. And an LLM works best here as one input into a deterministic, explainable engine, not as the final judge, especially when the attacker controls what it reads.
What's next for Guardian Loop
- Real phone integration through telephony providers (e.g., Twilio Media Streams) or an on-device app. The pipeline already accepts 8 kHz phone-quality audio.
- One-tap "join the call" so guardians can enter the conversation directly from the alert, plus SMS and push notifications.
- Scam Gym: practice calls against an AI scammer, scored by the same engine, with a debrief to help seniors recognize tactics before a real call.
- Persistent storage (MongoDB) and multi-guardian, multi-household accounts.
- Broader evaluation: our A/B test covered one scripted scam, so next we want a larger labeled dataset, precision/recall reporting, false-alarm testing on legitimate calls, and more languages.
Sources: FBI IC3 2025 Internet Crime Report (ic3.gov); FTC, Protecting Older Consumers 2024–2025 (Dec 2025); FTC imposter scam data release (June 2026); FTC Data Spotlight, "False alarm, real scam" (Aug 2025).
Built With
- audioworklet
- deepgram
- elderly
- express.js
- gemini
- github-actions
- google-genai
- javascript
- node.js
- npm
- react
- transparency
- typescript
- vite
- vitest
- vulnerable
- webrtc
- websockets
Log in or sign up for Devpost to join the conversation.