We will be undergoing planned maintenance on Oct 7th 6:00AM UTC / Oct 7th 2:00AM ET

Inspiration

As a solo developer building web applications, I faced a frustrating, daily reality: I don't have a team of QA engineers to test every button, form, and user flow before shipping.

Traditional end-to-end testing tools (Cypress, Selenium, Playwright test scripts) require engineers to manually write hundreds of lines of brittle test code that constantly break whenever a CSS selector or class changes. Even worse, generative AI models (like ChatGPT or Copilot) only read static text files — an AI cannot launch a real browser, click interactive buttons, or experience live runtime database deadlocks, network timeouts, and frozen loading spinners in a running application.

Apps usually don't fail because of bad code; they fail because real users click buttons in unexpected orders that solo developers never had the time or foresight to anticipate.

That inspired me to build BehaviorX: an autonomous, black-box testing platform that explores web applications from the outside in — exactly like a real human QA engineer — with zero access to source code. It maps live application states into an interactive state machine graph and captures production bugs with forensic evidence and 1-click visual replays.


What It Does

  • Autonomous Universal Crawler: Drives a real headless Chromium browser (via Playwright) across any live website on the internet (e.g., https://quotes.toscrape.com or https://news.ycombinator.com) or local web application, discovering routes, forms, and interactive elements.
  • Behavioral State Machine Mapping: Rather than a flat list of URLs, BehaviorX dynamically constructs an interactive directed state graph (powered by @xyflow/react) representing pages as state nodes and user actions as transitions, with a collapsible aerial overview MiniMap.
  • Cardinal Production Bug Detection: Automatically detects and flags critical failures:
    • 🔴 Client Runtime Crash: Uncaught JavaScript runtime exceptions (TypeError).
    • 🔴 Server Error (HTTP 500): Backend API crashes and transaction deadlocks with full request/response telemetry.
    • 🟠 Dead End State: Isolated views with zero outbound navigation links where users get stranded.
  • Observed Evidence Drawer: Separates human-readable explanations from forensic runtime facts, isolating the exact endpoint URL, HTTP status, DOM selector, JSON request payload, and stack trace.
  • Hero Feature — ⚡ 1-Click Deterministic Visual Replay: Spawns a fresh Playwright browser session, re-executes the exact user interaction breadcrumbs, captures live viewport snapshots step-by-step, and streams them into an interactive timeline player.
  • Embedded Dummy Benchmark App (NovaStore): Includes a dummy e-commerce hardware store (/demo-app) with 3 deliberate, deterministic production failure scenarios to benchmark the crawler out of the box.

How I Built It

  • Framework: Built with Next.js 14 (App Router) and TypeScript for type-safe full-stack performance.
  • Headless Browser Orchestration: Powered by Playwright (Chromium) running server-side in API route handlers to perform deep black-box DOM inspection over the Chrome DevTools Protocol (CDP).
  • Structural DOM State Hasher: Designed a custom fingerprinting engine (lib/crawler/state-hasher.ts) that extracts semantic landmarks, page titles, and interactive element densities to deterministically deduplicate and identify unique application states without source code.
  • Visual State Canvas: Implemented interactive directed graph visualization using @xyflow/react with custom dark-mode nodes and a collapsible aerial overview MiniMap.
  • Styling & Telemetry: Styled using Tailwind CSS with an industrial developer console aesthetic (zinc-950, hairline borders) and a real-time monospace streaming activity terminal.

Challenges I Ran Into

  1. State Deduplication in Dynamic SPAs: Modern single-page applications often update the DOM without reloading or changing the URL. Traditional web scrapers get stuck or create thousands of duplicate states. I engineered a DOM landmark hashing algorithm that strips volatile tokens and hashes structural headings and form action targets to accurately identify distinct states.
  2. Replay Timing Synchronization: In asynchronous form submissions, replaying actions too quickly resulted in locator timeouts. I implemented DOM readiness synchronization to ensure each replay step waits for navigation or DOM mutations before executing the next action.
  3. Handling Silent UI Deadlocks: Some bugs (like the checkout HTTP 500 error) cause UI buttons to spin indefinitely without rendering an error banner. I attached event listeners directly to the network layer (page.on('response')) to detect server errors even when the frontend fails to display an alert.

Accomplishments That I'm Proud Of

  • True Universal Crawling: Proved that BehaviorX is not hardcoded to a mock site by successfully crawling real, live internet websites like https://quotes.toscrape.com and https://news.ycombinator.com in under 7 seconds!
  • 1-Click Deterministic Replay: Engineered a real step-by-step browser replay engine that streams actual viewport snapshots rather than canned mockup animations.
  • Complete Solo Execution: Designed, built, and shipped a full developer tool with zero build errors and 15 clean Git commits during the hackathon.

What I Learned (Learning & Growth)

As a solo developer participating in First Commit, this project was a huge milestone in my engineering growth:

  • Headless Browser Protocols: Deepened my understanding of how browser automation communicates via the Chrome DevTools Protocol (CDP) to track console events, page crashes, and network requests.
  • State Machine Modeling: Learned how to model complex asynchronous user journeys as formal directed graphs with deterministic state transitions.
  • Forensic Tooling: Discovered the power of separating raw runtime evidence (payloads, selectors, status codes) from diagnostic explanations in developer tooling.
  • AI as a Pair Programmer: Learned how to effectively use AI tools as a learning aid to brainstorm heuristics and debug timing race conditions while driving the core architecture myself.

🤖 AI Usage Disclosure (Rules 5 & 7 Compliance)

In accordance with Hackathon Rules 5 and 7, AI assistance was used as an intelligent learning aid during the development of BehaviorX:

  • Assistance: Brainstorming behavioral exploration heuristics, scaffolding initial TypeScript interfaces, and assisting with debugging Playwright timing race conditions.
  • Human Authorship: All architectural decisions, state-machine modeling, and verification were directed and implemented by myself as a solo developer.

What's Next for BehaviorX

  • GitHub Actions / CI/CD Integration: Automatically running BehaviorX on pull requests to catch regression deadlocks before deployment.
  • Fuzz Testing & Edge-Case Mutation: Generating automated generative payload fuzzing to test form inputs with random boundary values.
  • Multi-Viewport Responsive Exploration: Testing mobile, tablet, and desktop viewports concurrently to catch responsive layout breaks.

Built With

  • black-box-testing
  • chromium
  • crawler
  • developer-tools
  • devtool
  • e2e-testing
  • headless-browser
  • javascript
  • next.js
  • node.js
  • playwright
  • qa-automation
  • react
  • react-flow
  • software-testing
  • state-machine
  • tailwind-css
  • testing
  • typescript
  • web-scraping
  • xyflow
Share this project:

Updates

Submission history