Add test_function to execution tests - #24
Merged
Merged
Conversation
Adds an optional test_function schema field holding the PHP signature of the entry point the verifier calls. The harness renders it to the model as a 'Define this function:' line and derives an unscored runtime function_exists assertion from the name, so the gateway function is communicated and diagnosable without earning score or requiring a static pattern. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Moves every gateway function signature out of prompts and static_checks into test_function across all 28 execution suites (124 tests migrated, 26 without a gateway function unchanged). Prompts now describe the task; test_function owns the entry-point contract, including previously undocumented signatures the verifier calls with arguments. Adds suite lints enforcing the contract: assertions invoking a wpbp_* function require a matching test_function, and the name may not appear in the prompt or static patterns. All 150 reference solutions pass --check-reference-solution. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
The following accounts have interacted with this PR and/or linked issues. I will continue to update these lists as activity occurs. You can also manually ask me to refresh this list by adding the If you're merging code through a pull request on GitHub, copy and paste the following into the bottom of the merge commit message. To understand the WordPress project's expectations around crediting contributors, please review the Contributor Attribution page in the Core Handbook. |
# Conflicts: # python/tests/test_reference_solution_mode.py # python/wp_bench/core.py
Contributor
|
verified live, all 150 ref solutions pass at 1.0, function_exists check stays weight-0 like you said, and missing-gateway drops to 0 with a clean diagnostic. 🙌 merging |
This was referenced Jul 10, 2026
lezama
added a commit
that referenced
this pull request
Jul 17, 2026
…its (#43) ## What Hardens the 8 execution tests that the #41 exploit audit (`--check-exploits`) flagged as exploitable: a zero-effort constant return (`return true;`, `return 1;`, `return false;`) satisfied their runtime assertions. All 8 shared the same root cause: the assertion graded the function's return value against a single fixed expectation instead of observing WordPress state, so it could not tell a real implementation from a hard-coded answer. ## How each test is hardened Every assertion now verifies behavior a constant cannot fake: | Test | Old weakness | New verification | |---|---|---| | `e-ai-client-001` | `return false` passed | Counting spy on the `wp_supports_ai` filter proves the support state is actually read during the call, plus the existing re-enabled-after check | | `e-ai-client-005` | `return false` passed | Late spy (`PHP_INT_MAX`) on `wp_ai_client_prevent_prompt` proves prevention is observed **active during** the support check, and inactive after | | `e-comments-users-004` | `return 1` passed | Second fixture post with zero comments; expected counts become `{2, 0}`, so no single constant passes both calls | | `e-connectors-api-001` | `return true` passed | Helper is called twice: with Akismet registered (expects `true`) and temporarily unregistered via `WP_Connector_Registry::unregister()` (expects `false`), then restored | | `e-post-types-taxonomy-001` | assertion returned the function's own `$ok` | Inspects `get_post_type_object( 'wpbp_book' )`: exists, `public`, `show_in_rest` | | `e-post-types-taxonomy-002` | same | Inspects `get_taxonomy( 'wpbp_genre' )`: exists, `hierarchical`, attached to `post` | | `e-post-types-taxonomy-003` | assertion was literally `return f();` | Fixtures (`wpbp_movie`, `wpbp_topic`) move to runtime `setup`/`teardown`; assertion checks the actual attachment with `is_object_in_taxonomy()` | | `e-roles-caps-002` | `return true` passed | Spy on `updated_option` for the roles option observes `wpbp_temp_cap` actually being **added** to Subscriber before the final not-present check | ## Design decisions - **Slugs moved into prompts.** Models only see `prompt` + `requirements` + `test_function` (never `expected_behavior`), so state-based assertions on `wpbp_book` / `wpbp_genre` / `wpbp_topic` / `wpbp_movie` require the prompts to name them. Done for the three post-types tests. - **Dropped the "Clean up registered types when possible" requirement** on `e-post-types-taxonomy-001/002`. A post-call state inspection requires the registration to still exist when the assertion runs; cleanup is now the assertion's job (as the reference solutions already assumed). - **`e-connectors-api-001` keeps its contract** (no-arg helper, `akismet` required pattern, `WP_Connector_Registry` forbidden for the model). The registry-toggle happens in assertion code, which static checks don't apply to. A lazy-init spy on `wp_connectors_init` was tried first but the registry initializes during `init`, before assertions run. - **`e-connectors-api-002`-style negative cases** were preferred over spies wherever the API allowed it; spies are only used where the observable is a filter/option side effect (`ai-client`, `roles-caps`). - Each touched suite file bumps `2.1.0` → `2.2.0`, same convention as #24. ## Verification (live WP 7.0 runtime) Both directions of the audit loop, against the real wp-env grader: - `--check-reference-solution` on the 8 tests: **8/8 pass** (correctness 1.0) - `--check-exploits` on the 8 tests: **0/8 exploitable** - Full pytest suite: 151 passed; dataset validation green - `--check-exploits` on the full suite: **0/140 auditable tests exploitable** (was 8/140 at the #41 baseline). Full end-to-end run: 1120 exploit-candidate executions, each against a freshly reset WordPress, ~1.5h wall clock. The 10 non-auditable tests have no `test_function` or plugin artifact for the generic battery to target, unchanged from the baseline. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This was referenced Aug 5, 2026
Closed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Execution tests need a gateway function the verifier can call, but that function is harness scaffolding — not WordPress knowledge worth scoring, and not something the model should have to guess. Several tests invoked functions with arguments or return shapes the prompt never described (e.g. `e-queries-009`'s `bool $previous` flag, `e-post-types-taxonomy-005`'s `$name` argument), making them unfair. This PR makes the entry-point contract explicit, machine-checked, and unscored.
What changed
Harness (first commit):
Datasets (second commit):
Verification
🤖 Generated with Claude Code