ShadowFrog gives coding agents a shadow knowledge base for any codebase: a file-backed memory of tacit codebase knowledge learned from code reading, experiments, and conversations with you.
A shadow mirrors your source tree under .shadow/, storing discoveries in
symbol-organized Markdown files. Lookup is index-free: agents follow source
paths and file::symbol references rather than a vector index or embedding
service.
It records knowledge that is hard to recover from source alone: which refactor breaks downstream callers, which invariant the tests never exercise, or which "obvious" cleanup removes a production workaround. The code tells you what runs. The shadow tells future agents what has been learned about how it behaves.
Read the launch blog post: Shadow-Frog: Coding Agents that Dream and Discover.
Install per repository, not globally. You need Git, Python 3, either GitHub Copilot CLI or Claude Code, and a Git repository as the target project. Hooks are limited to projects you explicitly opt into.
From a checkout of ShadowFrog, choose one agent:
cd /path/to/ShadowFrog
./install.sh --project /path/to/your-repo # Copilot CLI (default)
# ./install.sh --agent claude --project /path/to/your-repo # Claude CodeOn Windows, use the PowerShell installer, which does not require Bash:
cd C:\path\to\ShadowFrog
.\install.ps1 -Project C:\path\to\your-repo # Copilot CLI (default)
# .\install.ps1 -Agent claude -Project C:\path\to\your-repo # Claude CodeFor a shared setup, commit and push the installed files in the target repo. The installer prints the exact staging paths for your selected components:
cd /path/to/your-repo
# Stage the files for the agent you installed:
git add .github/skills/ .github/hooks/ .github/copilot-instructions.md
# git add .claude/skills/ .claude/hooks/ .claude/settings.json CLAUDE.md
git commit -m "Add ShadowFrog skills, hooks, and context"
git pushLocal-only use does not require a writable remote; Dream does.
To start collecting shadow knowledge, open the target repo in your agent session and run the following. Nap-only ideation can skip this step.
/shadow-frog-init
This creates symbol-organized templates in .shadow/. Choose how to store it:
| Mode | Effect |
|---|---|
| Committed | Team-shared knowledge; required for Dream |
| Gitignored | Local-only knowledge; update, meditate, viewer, and Nap remain available, but Dream is disabled |
If you chose committed storage, commit and push .shadow/ after initialization.
For installed paths and optional components, see Installation options.
Use the skill that matches your goal. Each link contains its full workflow, helper commands, and format definitions.
| Command | When to use |
|---|---|
/shadow-frog |
Consult relevant knowledge before editing or investigating code |
/shadow-frog-init |
Create the shadow once per repo |
/shadow-frog-update |
Refresh after code changes and capture session insights |
/shadow-frog-dream |
Run autonomous experiments while you're away |
/shadow-frog-nap |
Generate reviewed feature-task briefs without implementing them |
/shadow-frog-meditate |
Merge duplicates and resolve conflicting discoveries |
/shadow-frog-viewer |
Browse, search, inspect lineage, and audit structural integrity |
As you work, the agent captures your code context as source: user and
collaborative findings as source: interaction. After commits, the pre-tool
hook can detect a shadow behind HEAD and remind the agent to run
/shadow-frog-update; the hook does not run the update itself. Meditate
consolidates accumulated knowledge and escalates unresolved conflicts to you.
For example, use Viewer to find relevant knowledge or audit its structure:
/shadow-frog-viewer --search "auth"
/shadow-frog-viewer --top src/auth.py
/shadow-frog-viewer --check-invariants
The Viewer reference also covers summaries, recent discoveries, label filters, preferences, and interactive dream-lineage HTML.
| Dream | Nap | |
|---|---|---|
| Goal | Learn through implemented experiments | Develop source-grounded feature/task proposals |
| Output | Runnable experiment branches and discoveries | Reviewed task briefs and a persistent proposal tree |
| Continuation | Inherit code and shadow from an ancestor branch | Revise hypothetical designs over a pinned code baseline |
| Implementation | Write and run real code in isolated worktrees | No feature code or prototypes; optional probes inspect existing behavior |
| Prerequisites | Initialized, git-tracked shadow and writable remote | A Git repository with a commit; no remote or initialized shadow required |
Experiments persist as dream/<namespace>/<id> branches. Future dreams can
continue a previous experiment, inheriting its code and shadow rather than
sibling branches. Reconciliation accumulates discoveries and experiment
reports on the default branch.
Before running Dream, commit and push .shadow/ and configure a remote
that permits pushing dream/... branches and reconciled shadow updates.
Gitignored shadows cannot use Dream. Experiment code is not merged
automatically; adopting it into the project is a manual curation step.
See the Dream workflow for execution, tooling snapshots, reconciliation, and safe cleanup.
Nap grows a resumable proposal tree and submits shortlisted paths to a strong independent judge. Exported tasks require a current accepted judgment and describe the complete change from a real code baseline, not assumed parent APIs. Planning approval is not runtime validation. The host runs the judge; the Python helper manages records and cannot authenticate reviewer identities.
Default limits are 7 recorded nodes, 2 probes, 2 judge batches, and 2 selected
tasks, with no default depth cap. These configurable ceilings are not quotas:
recorded rejections and failed probes count. They do not cap actual API spending.
Keep proposals outside .shadow/, or in an initialized .shadow/_meta/naps/;
they are not verified discoveries.
Exports can be detailed planning briefs or concise implementation handoffs, with the same active requirements. Binding constraints are separate from design suggestions; nonblocking implementation risks are separate from questions that prevent planning approval.
See the Nap workflow and helper reference for tree operations, review receipts, and exports.
Dream and Nap default to mode=broad. Use mode=coherent to require a
meaningful connection along each parent-child edge:
/shadow-frog-dream mode=coherent
/shadow-frog-nap mode=coherent
Children can extend, integrate, challenge, replace, simplify, or offer an alternative to their parent. Each has its own goal; siblings can pursue different directions, including on the same files. There is no fixed tree-wide goal or diversity quota.
Dream descendants wait for their implemented parent; coherent branches and their ancestors are retained as reproducible baselines until explicit curation. Nap compounds ideas, not implemented APIs. A combined task follows a root-to-leaf path and preserves the final active requirements, rather than stacking unrelated siblings or requiring both discarded and replacement designs. Combining sibling work requires an explicit integration experiment or proposal. Structural validation alone cannot establish semantic coherence or feasibility.
For example, knowledge about src/auth.py lives at .shadow/src/auth.py.md;
locations such as src/auth.py::login identify the relevant symbol.
Cross-file discoveries live once in _cross/, with links from the involved
per-file shadows.
your-repo/
src/auth.py
.shadow/
src/auth.py.md file- and symbol-level discoveries
_cross/ cross-cutting discoveries
_prefs.md project-wide preferences
_dreams/ experiment reports, manifests, and patches
_index.md file inventory and counts
_meta/state.json update state
.shadowignore gitignore-style exclusions
Store behavioral discoveries, not API descriptions or chat transcripts. For example:
Agent exploration:
- authenticate_user() silently returns None on expired tokens
instead of raising. 3 of 7 callers don't check the return value.
_(verified, source: exploration, labels: [bug])_User knowledge:
- The retry logic here took 3 iterations to get right -- it handles
a subtle race condition during rolling deployments. Do not simplify.
_(verified, source: user)_Collaborative work:
- While debugging issue #42, discovered that process_batch() silently
drops items exceeding 1MB -- logged at DEBUG level only.
_(verified, source: interaction)_| Property | Values | Meaning |
|---|---|---|
| Status | verified / uncertain / refuted |
Has the claim been confirmed? |
| Source | exploration / user / interaction |
Where did this knowledge come from? |
| Labels | bug, performance, security, feature-gap, tech-debt |
Optional; marks actionable discoveries |
| Rank | Source | Trust |
|---|---|---|
| 1 | source: user |
Highest; human stated it. Always verified. |
| 2 | source: interaction |
Emerged from collaborative work. Always verified. |
| 3 | verified, source: exploration |
Agent confirmed via code analysis or tests. |
| 4 | uncertain |
Plausible but unconfirmed. |
| 5 | refuted |
Known wrong; skip. |
See the worked coupon example for a real shadow, and the core skill for the canonical formats and reference rules.
The installer copies readable skill instructions, their helper scripts, hooks, and agent context. It targets one agent's conventions at a time:
| Agent | Skills | Hooks | Context |
|---|---|---|---|
copilot (default) |
.github/skills/ |
.github/hooks/hooks.json |
.github/copilot-instructions.md |
claude |
.claude/skills/ |
.claude/settings.json |
CLAUDE.md |
Use --no-hooks or --no-context (-NoHooks / -NoContext in PowerShell) to
skip individual components. On Windows, py -3 can be used when invoking
Python helpers. Full helper usage is in the linked skill references above.
| Path | Purpose |
|---|---|
skills/ |
The seven skills and their helper scripts |
hook-templates/ |
Agent hook configs and shared scripts |
examples/coupon-demo/ |
Worked example with a real .shadow/ |
eval/ |
Evaluation methodology and results dashboard |
tests/ |
Installer, hook, and helper regression coverage |
Install development dependencies and run the suite:
pip install -r requirements-dev.txt
python3 -m pytestTests use temporary shadow trees and Git repositories. Their layout mirrors
the source: tests/skills/shadow_frog_viewer/ covers
skills/shadow-frog-viewer/, for example.
ShadowFrog is a research project. Before using it, please review our Responsible AI transparency note, which covers intended uses, out-of-scope uses, evaluation, limitations, and best practices.
This project is licensed under the MIT License. See the LICENSE file for details.
Built by the Froggy team
