Skip to main content
ROI Scale AI logoROI Scale AI
Business
Technology & Telecom
arrow_forward
Financial Services
arrow_forward
Healthcare
arrow_forward
Retail & E-Commerce
arrow_forward
Education
arrow_forward
Energy & Utilities
arrow_forward
Media & Entertainment
arrow_forward
Manufacturing & Industrial
arrow_forward
Real Estate & Construction
arrow_forward
Government & Public Sector
arrow_forward
Professional Services
arrow_forward
Transport and Logistics
arrow_forward
View all in Business arrow_forward
Technology
Models & Benchmarks
arrow_forward
AI Engineering
arrow_forward
Harness Engineering
arrow_forward
Data Strategy
arrow_forward
AI Security & Governance
arrow_forward
Libraries & Frameworks
arrow_forward
AI for Developers
arrow_forward
Research & Papers
arrow_forward
View all in Technology arrow_forward
Marketplace
Blueprints
arrow_forward
Proof Packs
arrow_forward
View all in Marketplace arrow_forward
Contribute
How-Tos
arrow_forward
Business RoadMap
arrow_forward
Tech RoadMap
arrow_forward
View all in Contribute arrow_forward
search
person_outlineSign In
Categories
BusinessTechnology & TelecomFinancial ServicesHealthcareRetail & E-CommerceEducationEnergy & UtilitiesMedia & EntertainmentManufacturing & IndustrialReal Estate & ConstructionGovernment & Public SectorProfessional ServicesTransport and Logistics
TechnologyModels & BenchmarksAI EngineeringHarness EngineeringData StrategyAI Security & GovernanceLibraries & FrameworksAI for DevelopersResearch & Papers
MarketplaceBlueprintsProof Packs
ContributeHow-TosBusiness RoadMapTech RoadMap
searchSearchhomeHome
Community
person_outlineSign In / Join
Home/Marketplace/Proof Packs
Browser-Use Frameworks Compared (Playwright-MCP vs Browser-Use vs Computer-Use)
zoom_in
Browser-Use Frameworks Compared (Playwright-MCP vs Browser-Use vs Computer-Use)
Browser-Use Frameworks Compared (Playwright-MCP vs Browser-Use vs Computer-Use)
benchmark

Browser-Use Frameworks Compared (Playwright-MCP vs Browser-Use vs Computer-Use)

by Admin

Free

A runnable harness that puts Playwright-MCP, Browser-Use, and Anthropic Computer-Use through the same 30 real workflows — so you get the success rates, partial-credit trap, and failure-mode taxonomy the README demos will never show you, plus the tiered router that actually hit 71% automated in production.


Browser control is the capability that unlocks agents actually doing things on the web — booking, submitting, filling, navigating. The numbers people post online are almost always cherry-picked demos. This blueprint is the companion code to Browser-Use Frameworks Compared: Playwright-MCP, Browser-Use, and Computer-Use Against 30 Real Workflows — a 30-workflow catalog, a three-valued scoring layer, fixture sites that isolate one failure mode each, and the tiered router the author actually deployed.


What's inside:

·        30 workflows across 4 categories — 10 transactional, 8 extraction, 7 migration/bulk, 5 form completion — pinned in config/tasks.yaml

·        Framework adapters for Playwright-MCP (DOM), Browser-Use (GPT-4o agent loop), and Computer-Use (pixel-level Chrome in Docker), with the article's published rates baked into config

·        Partial-credit scoring (full / partial / fail) so "got to checkout, never clicked confirm" is visible instead of silently counted as a win

·        Fixture sites that isolate the failure modes: SPA data-testid regeneration, soft paywall, odd multi-page forms, bulk tables, checkout interstitials, and MFA

·        A tiered router — Browser-Use first (cheapest/fastest), Playwright-MCP on first failure, human after the second

·        Dry-run mode: mock adapters reproduce the published rates with zero API keys, safe for CI or a local demo

What the numbers actually said: Playwright-MCP won on strict success (52%) but still broke on SPA re-renders; Browser-Use was fastest and cheapest (1:40, $0.06/task) but over-clicked (18 actions vs 11) and only solved 20% of forms; Computer-Use won on odd forms (60%) and lost on time (4:20 mean, $0.42/task). Add 15–22 points to all three if you count one-step-short as success. All three fail on MFA. The deployed tiered system hit 71% fully automated / 29% human-assisted.

Get the code: github.com/mskrado/poc/browser-agent-benchmark — MIT licensed, pip install -e ".[dev]", python demo.py --scenario all.

Read the full breakdown: Browser-Use Frameworks Compared: Playwright-MCP, Browser-Use, and Computer-Use Against 30 Real Workflows


Loading ratings...

Comments (0)

Join the conversation!

Loading comments...

Support

  • Contact Us
  • FAQ
  • Privacy
  • Terms

About

  • Mission
  • Editorial

© 2026 ROI Scale AI. All rights reserved.

Powered by Publishi.ai