Commit Graph
66 Commits
Author SHA1 Message Date
SpencerWhitehead c54f6c2c8a update readme 2026-07-22 07:43:57 -07:00
Hussein Mozannar 08d0363a35 updates 2026-07-17 13:04:48 -07:00
Hussein Mozannar 2322f2f674 recover 2026-07-17 13:03:07 -07:00
Hussein Mozannar 879afdd82c more cleanup 2026-07-17 12:59:27 -07:00
Hussein Mozannar 5c8247a5f2 minor 2026-07-17 12:52:27 -07:00
Hussein Mozannar eda8e982d8 clean up 1 2026-07-17 12:31:25 -07:00
Hussein Mozannar 8e33bdec69 comments 2026-07-17 12:18:04 -07:00
Hussein Mozannar 522a9ccbae remove 2026-07-17 12:14:04 -07:00
Hussein Mozannar d58a16cb6a stage1 2026-07-17 11:46:48 -07:00
Hussein Mozannar 16aa6c0007 readme update 2026-07-17 10:32:14 -07:00
Hussein MozannarandClaude Fable 5 5274e1ee06 Delete empty .gitmodules
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-17 09:58:20 -07:00
Hussein MozannarandClaude Fable 5 ca3da813b9 Remove unused autogen submodule
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-17 09:58:06 -07:00
Corby Rosset 9b8af39c87 Port CP-aware verifier + error taxonomy (#77) 2026-06-16 19:27:57 -04:00
corbyandClaude Opus 4.8 c529b48511 llm_helpers: unwrap CreateResult.content defensively (fix Step 9b on gateway)
llm_call_expect_json did `result.content.content`, assuming a nested object.
On the OpenAI/phyagi-gateway path CreateResult.content is already a plain
str, so this raised "'str' object has no attribute 'content'" on every
attempt — Step 9b (trajectory-informed task verification) failed all retries
and silently fell back to a default. The bug was masked because the test only
asserts the result keys exist, not that the step ran.

Mirror task_classification's helper: unwrap .content only when present. Step
9b now runs clean against the gateway (no retries, real verdict).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-29 12:52:36 -07:00
corbyandClaude Opus 4.8 19224e97da graceful_client: guarantee create() terminates; fix infinite retry hang
A live verifier test wedged for ~19h: a judge endpoint returning HTTP 401
(mapped by the openai SDK to AuthenticationError) hit the AuthenticationError
branch, which neither blocklisted the endpoint nor consumed the `tries`
budget — so next_client() kept handing back the same failing endpoint and
the loop spun forever, opening connections the whole time.

- AuthenticationError and the check-access-response-enc branch now decrement
  `tries` (and AuthenticationError backs off 1s) so a persistent auth/access
  failure on a small pool can't loop forever.
- Add a hard `max_total_attempts` cap (default max_retries + 2*n_endpoints)
  as a backstop: create() always terminates regardless of which branch fires,
  including blocklisting paths that intentionally don't spend the budget.
- Exhaustion error now reports total_attempts + both caps.
- Tests: 3 regression cases (persistent-auth termination across 1 and 3
  endpoints, all-endpoints-blocklisted termination), each guarded by a 30s
  asyncio.wait_for so any future regression fails loudly instead of hanging.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-29 12:52:36 -07:00
corbyandClaude Opus 4.7 13f9e77d17 Port CP-aware verifier + error taxonomy from agento_next
Brings the rubric_agent package up to parity with agento_next/main
(post-#1071 architecture), adapted to webeval:

- Adds `CriticalPointAgent` (Step -1) that classifies the task against
  a critical-point taxonomy (`critical_point_types.yaml`) and threads
  the result through rubric generation, action-only scoring, outcome
  verification, and a new CP-violation check.
- Adds `VerifierAgent` (Steps 9a/9b/10): first-point-of-failure
  analysis with the error taxonomy, trajectory-informed task
  verification, and unified task verification. Step 11 (synthetic
  human-voice feedback) is intentionally dropped — not needed in webeval.
- Mirrors upstream #1071's DRY refactor: extracts shared helpers
  (`format_action_history`, `call_llm`, `encode_image_b64`,
  `get_init_url_context`, `build_scored_rubric_summary`,
  `build_all_screenshot_evidence_text`) into `formatting.py` so
  `MMRubricAgent` and `VerifierAgent` don't duplicate them. Also
  removes the lazy `MMRubricAgent._format_action_history` import
  from `critical_point_classifier.py`.
- Picks up upstream #889 `run_command` support: adds
  `StepSummary.tool_output`, populates it from post-action
  `ToolOutput` observations, renders a `Command Output:` line in the
  action history, and teaches the prompts to treat that output as
  ground truth (so an unchanged desktop after `run_command` is not
  read as failure) while sanity-checking that the command isn't a
  fake (`echo "success"`).
- Adds missing runtime deps (`imagehash`, `jinja2`) to
  `webeval/pyproject.toml` — both are imported by the new modules.

Architecture note: the branch keeps `MMRubricAgent` and `VerifierAgent`
as independent agents orchestrated by the caller, rather than upstream
#1071's compose pattern. This is the same direction (decoupling) and
goes a step further by also stripping Step 11. `verify_trajectories.py`
drives them in sequence.

All 28 webeval unit tests pass; the live-LLM end-to-end test
(`test_verify_trajectories_live_llm`, opt-in via `FARA_VERIFY_LIVE_TEST=1`)
asserts the new CP-aware fields (`cp_type_used`, `cp_violation`,
`error_taxonomy.first_point_of_failure.failure_points[].error_code`,
Steps 9b/10 `is_ambiguous`/`is_invalid`) hit the score file.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-28 13:49:05 -07:00
Hussein Mozannar 796ae4f2bf Update README (#76) 2026-05-21 17:27:42 -04:00
Hussein Mozannar 29bc9c33a0 Fix formatting in update entry for Fara 1.5 2026-05-21 17:27:30 -04:00
Hussein Mozannar 1a7ad8384d Update README
update readme for fara 1.5 coming soon
2026-05-21 17:26:57 -04:00
Corby Rosset 8d823d72c3 docs: WebTailBench V1↔V2 diff page + README refresh note (#74) 2026-05-12 18:22:24 -04:00
corbyandClaude Opus 4.7 80d96b1ffc README: rename split test-v2 -> test_v2
HF datasets metadata does not allow '-' in split names. Match the
corrected split name on microsoft/WebTailBench.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-12 15:16:43 -07:00
corbyandClaude Opus 4.7 406b26d4a3 docs: add WebTailBench V1↔V2 diff page; note refresh in README
V2 refreshes 609 WebTailBench tasks and their precomputed rubrics
(V1 was calendar-bound through Nov 2025). New side-by-side diff page
under docs/ shows per-task changes in task_summary plus a collapsible
unified diff of the precomputed_rubric JSON for each task.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-12 15:10:51 -07:00
Corby Rosset 44908264c8 Fix NameError from orphaned TrajectoryDiagnosticsResult reference (#68) 2026-04-23 01:42:24 -04:00
corbyandClaude Opus 4.7 ea3ce6eac4 Drop leftover TrajectoryDiagnosticsResult reference
Commit acc1404 removed the TrajectoryDiagnosticsResult class but left
it inside the VerificationResultEvent discriminated Union, which made
`from webeval.rubric_agent.data_point import *` raise NameError at
import time and broke every test that transitively imports the module
(21 failures / errors across test_rubric_agent_imports,
test_shared_data_adapter, test_webtailbench_dataset,
test_verify_trajectories). Removing the orphan reference restores the
test suite to 27 passed / 2 skipped.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-22 22:41:39 -07:00
Corby Rosset cde2ebcd6d Corby/universal verifier (#66) 2026-04-23 01:39:12 -04:00
Corby Rosset acc14047f0 Remove TrajectoryDiagnosticsResult class
Removed the TrajectoryDiagnosticsResult class and its attributes from data_point.py.
2026-04-23 01:38:28 -04:00
Corby Rosset b0030f7abe Update __init__.py 2026-04-23 01:33:57 -04:00
corbyandClaude Opus 4.7 7ea1e441c4 Add example_trajectory + expand webeval trajectory-format docs
Drops the cuaverifierbench/ build script (now hosted alongside the
dataset on HuggingFace) and checks in a full WebTailBench trajectory
under webeval/data/example_trajectory/ (web_surfer.log, final_answer,
screenshots, core.log, times.json, task_data.json, rubric score file).
webeval/README.md now walks through each file against that concrete
example and corrects several field-level inaccuracies (times.json keys,
emitted event types, webtailbench score payload, auto-0 vs excluded
semantics). Adds test_trajectory_loading and test_verify_trajectories
coverage; repairs HuggingFace dataset URLs and doubled /path/to paths
in the root README.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-22 22:30:04 -07:00
corbyandClaude Opus 4.7 8cd48183e0 Add CUAVerifierBench dataset (build script + README + repo references)
CUAVerifierBench is a human-annotated benchmark for evaluating CUA
verifiers (judges that score agent trajectories). Released to
huggingface.co/datasets/microsoft/CUAVerifierBench as two configs
(trajectories + annotations) joinable on task_id, with two splits:

- fara7b_om2w_browserbase: 106 Fara-7B Online-Mind2Web/Browserbase
  trajectories x ~2 reviewers (UV-blind + UV-informed labels)
- internal: 154 trajectories from a heldout aurora-v2 task suite,
  single reviewer per task (UV-blind only)

This commit adds:
- cuaverifierbench/build_dataset.py — builder script
- cuaverifierbench/README.md — dataset card mirrored to HF
- README.md — new badge, Updates entry (2026-04-19), and a
  CUAVerifierBench section after the WebTailBench results table

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-18 20:01:57 -07:00
corbyandClaude Opus 4.7 9f14b6e340 Universal Verifier (MMRubricAgent) + WebTailBench, autogen-free clients
Three changes in one PR:

1. Remove webeval's dependency on ``autogen-core`` / ``autogen-ext``.
   All chat completion clients, message types, and the graceful-retry
   layer now live under ``webeval/src/webeval/oai_clients/`` —
   self-contained wrappers around openai / azure-identity. Install no
   longer needs the autogen submodule; just ``pip install -e .[vllm]``
   then ``cd webeval; pip install -e .``.

2. Incorporate the initial (now-stale) WebTailBench benchmark into the
   codebase. ``webeval/src/webeval/benchmarks/webtailbench/`` +
   ``webeval/scripts/webtailbench.py``. Loader auto-downloads
   ``WebTailBench-v1-rubrics.tsv`` from
   ``huggingface.co/datasets/microsoft/WebTailBench`` and threads each
   task's published ``precomputed_rubric`` through to the verifier so
   rubrics never get regenerated.

3. Release the Universal Verifier (``MMRubricAgent``) as the official
   judge for WebTailBench. Multimodal, rubric-grounded, two-model
   ensemble (``gpt-5.2`` + ``o4-mini``) with per-criterion scoring,
   outcome verification, ambiguity / invalid-task classification, and
   first-point-of-failure analysis. ``webeval/scripts/verify_trajectories.py``
   is a stand-alone parallel runner that re-scores any directory of
   webeval-shaped trajectories without touching the solver.

Documentation: repo-root README ``Updates`` section + Reproducibility
CLI block; ``webeval/README.md`` documents the Trajectory / FinalAnswer
schema, the ``<no_answer>`` semantics, and per-benchmark score-file
shape.

Tests: 18 passing, 1 skipped (opt-in HF download).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-18 17:15:46 -07:00
Corby Rosset 2c27710499 Add WebTailBench rubric comparison visualizer (#65) 2026-04-17 00:34:52 -04:00
corbyandClaude Opus 4.7 2ec67d7236 Add WebTailBench rubric comparison visualizer
Standalone HTML page that shows, for each WebTailBench task, the
rubric criteria produced by three different judge configurations
side by side:

  1. O4-Mini Rubric                    — historical baseline
  2. GPT-5 (v1)                        — original GPT-5 judge
  3. Universal Verifier Rubric (GPT-5.2) — current release

Tasks are grouped by WebTailBench benchmark category, with
incremental search across ids / summaries / criteria, and an
"all three rubrics only" toggle. The header links directly to the
microsoft/WebTailBench dataset on Hugging Face and to the
WebTailBench-v1-rubrics.tsv download so readers can grab the
underlying data.

Source layout:
  docs/webtailbench_rubric_comparison.html  (single self-contained file)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-16 21:12:47 -07:00
Spencer Whitehead 0aa8a48430 update README (#52) 2026-01-14 15:09:53 -05:00
Andrew Zhao e648c0e844 update README 2026-01-14 15:03:24 -05:00
Hussein Mozannar 6261cdb6fc Add citation for Fara-7B in README (#47) 2025-12-15 11:33:48 -05:00
Spencer Whitehead 76f38253e9 Add citation for Fara-7B in README
Updated citation information for Fara-7B and added a link to the associated paper.
2025-12-15 11:17:43 -05:00
Hussein Mozannar 3f3f1fd82c docs: fix typos and formatting in README (#44) 2025-12-10 12:08:47 -05:00
Mehul Jariwala 9d7efb0e11 Fix typos and enhance clarity in README.md
Corrected typos and improved clarity in README.
2025-12-10 10:52:24 +05:30
Hussein Mozannar 6cbb4f3694 Windows Support (#40) 2025-12-03 17:28:43 -05:00
Hussein Mozannar 82637b54df fixes 2025-12-02 14:01:05 -08:00
Corby Rosset 879dc9561a online mind2web works (#23) 2025-12-02 00:16:35 -05:00
corby a3e2750c6a Merge branch 'main' into corby/om2w 2025-12-01 21:13:43 -08:00
Hussein Mozannar 5ebe2590f8 Update README with Magentic-UI instructions (#37) 2025-11-28 20:49:34 -05:00
Hussein Mozannar 5745ea715f Enhance README with video demos for Magentic-UI
Added video demos reference for Magentic-UI integration.
2025-11-28 20:49:18 -05:00
Hussein Mozannar 2b81016788 Update README with Magentic-UI instructions
Added instructions for using Fara-7B with Magentic-UI and included a note about WSL2 for Windows users.
2025-11-28 20:48:25 -05:00
Hussein Mozannar f10b9319b3 Refine Fara-7B description in README (#31) 2025-11-28 14:30:48 -05:00
Hussein Mozannar ccdc3def6e Refine Fara-7B description in README
Removed redundancy in the description of Fara-7B's visual operation.
2025-11-28 14:30:31 -05:00
Hussein Mozannar 21469308d6 Update README with Windows and memory usage notes (#30) 2025-11-28 13:35:43 -05:00
Hussein Mozannar b555f0874b Update README with Windows and memory usage notes
Added notes for Windows users and memory command.
2025-11-28 13:34:46 -05:00
corby 01f8a17ffd forgot om2w score file type 2025-11-27 11:37:43 -08:00