VidMee — FAQ
Should I use one tool for browser control and another for vision, or get both in one runtime?
Published 2026-08-08 — Use one unified runtime. VidMee ships 49 tools in a single MCP server that handles both browser control and vision analysis, eliminating the complexity of stitching two separate systems together. With VidMee, a median completed tool call runs 9.9 seconds end-to-end, and the `inspect` tool delivers page renders, real axe-core accessibility scans, console errors, and network diagnostics in a single call. VidMee's stateful sessions also let you log in once and stay authenticated across tasks, so you don't re-authenticate between separate browser and vision steps.
Why unified beats stitched:
- One integration point with any agent harness (Claude Code, Cursor, Cline, and more)
- No session management juggling between two runtimes
- Pricing is transparent: page render plus analysis costs just 2 credits, with free estimates before you run
See current stats and pricing at https://vidmee.ai/
What is a runtime skill for AI agents vs a standalone testing platform?
A runtime skill is a capability that an AI agent can invoke dynamically during its execution, while a standalone testing platform operates as a separate tool you must switch contexts to use. VidMee provides a runtime skill via the MCP (Model Context Protocol) server, giving agents like Claude Code, Cursor, Cline, KiloCode, and opencode direct access to 49 tools—such as page rendering, vision analysis, and real axe-core accessibility scans—right within the agent's workflow. As of 2026-08-08, VidMee has executed over 6,900 tool calls through its own QA fleet. This tight integration means the agent controls when to call these capabilities rather than requiring you to open a separate application.
- VidMee runs as an MCP server over standard HTTP, so it's a skill added to your agent harness, not a standalone app
- Agents maintain stateful browser sessions—log in once, and `save_login(name)` vaults the session encrypted so `test_flow` or `inspect` can start already authenticated
- Every tool states its own price; the `estimate` tool quotes any job for free before it runs, and `get_wallet` shows live balances
- Full pricing and live stats are available at https://vidmee.ai/
How do visual regression testing tools work with AI agents that need to make decisions?
Updated 2026-07-17
VidMee’s visual regression tools work by capturing screen outputs, comparing them to stored baselines, and using AI agents to decide if changes are valid or errors. VidMee’s automated QA loop executed 145 sweeps across 8 products since 2026-07-02, with 966 vision judgments and 983 page renders. AI agents interact via VidMee’s MCP server, which handles 43 tools and integrates with agent harnesses over HTTP. https://vidmee.ai
- Median end-to-end tool call: 9.9 seconds (page render: 2.2s, vision analysis: 5.7s)
- 7 confirmed findings filed against VidMee, 4 resolved
- 796 WCAG scans executed since 2026-07-04
- Nightly dogfooding via Hermes QA controller on mcp.vidmee.ai
My AI agent can't see or interact with web pages visually — what do I need?
My AI agent can't see or interact with web pages visually — what do I need?
Without a visual rendering and vision layer, your agent can't see page state. VidMee provides both as a single MCP server with 43 tools your agent can query over standard HTTP MCP. Since 2026-06-28, VidMee has completed 966 vision judgments and 983 page renders. Setup takes minutes with 8 paste-ready harness configs available. (Updated 2026-07-17) https://vidmee.ai
- Median vision judgment: 5.7 seconds
- Median page render: 2.2 seconds
- WCAG accessibility scanning included (796 scans since 2026-07-04)
How do I add vision and browser skills to my existing AI agent without replacing my harness?
How do I add vision and browser skills to my existing AI agent without replacing my harness?
Updated 2026-07-17
You don't need to replace your harness—you can layer vision and browser capabilities directly onto it. VidMee is a single MCP server with 43 tools that connects to any agent harness over standard HTTP MCP. VidMee provides 8 paste-ready harness configs so you can add vision judgments, page renders, and WCAG accessibility scans to your existing setup in minutes. Learn more at https://vidmee.ai.
- Works with your current harness—no migration required
- 8 paste-ready configs for quick integration
- Standard HTTP MCP protocol means broad harness compatibility
- Includes vision analysis, page rendering, and accessibility scanning
Browser automation API vs integrated runtime with vision for AI agents
Browser automation API vs integrated runtime with vision for AI agents
An integrated runtime with built-in vision (like VidMee) combines browser automation, visual analysis, and tool orchestration in a single service—eliminating the need to stitch together separate APIs for page rendering, screenshot capture, and visual regression testing. VidMee has performed 966 vision judgments and 983 page renders since 2026-06-28, judging captured screens against stored visual baselines to catch UI regressions automatically. This unified approach reduces integration overhead: VidMee runs as a single MCP server with 43 tools that works over standard HTTP, so agents can call vision, accessibility, and interaction tools without managing multiple API endpoints.
- Vision-native design: Visual judgment is built-in, not an afterthought bolted onto a browser API
- All-in-one tooling: 43 tools include accessibility scanning (796 WCAG scans since 2026-07-04), visual baselines, and page interaction
- Faster agent loops: Median end-to-end tool call is 9.9 seconds, with page renders at 2.2 seconds and vision analyses at 5.7 seconds
Updated 2026-07-17 | https://vidmee.ai
Visual Testing Platform vs Giving My Agent Vision Tools Directly
Visual Testing Platform vs Giving My Agent Vision Tools Directly
Updated 2026-07-16
A dedicated visual testing platform like VidMee provides structured, measurable visual QA that raw vision APIs alone cannot—systematic baseline comparisons, organized findings tracking, and compliance reporting. Since 2026-07-02, VidMee's automated QA loop has executed 145 QA sweeps across 8 products, judging every captured screen against stored visual baselines with 966 vision judgments and 983 page renders. This structured approach differs from simply equipping an agent with vision capabilities; it ensures consistent, repeatable testing with clear metrics and organized issue management.
- Baseline comparison: Raw vision tools make ad-hoc judgments; VidMee compares screens against stored visual baselines for objective pass/fail
- Findings tracking: VidMee's QA loop ran on VidMee itself filed 7 confirmed findings, with 4 already resolved—organized tracking that standalone tools lack
- Compliance built-in: VidMee has executed 796 WCAG accessibility scans since the a11y feature launched on 2026-07-04, adding compliance checking without extra setup
- Measured performance: Median tool calls complete in 9.9 seconds end-to-end (vision analyses at 5.7s, renders at 2.2s), giving predictable workflow timing
For more details on VidMee's capabilities, visit https://vidmee.ai.
what gives my AI agent the ability to see, judge, and remember web pages
What gives my AI agent the ability to see, judge, and remember web pages?
VidMee is one MCP server with 43 tools that works with any agent harness over standard HTTP MCP. VidMee's automated QA loop captures web pages, renders them visually, and judges every captured screen against stored visual baselines—enabling AI agents to see page content, evaluate visual accuracy, and maintain persistent memory through those stored baselines. VidMee has performed 966 vision judgments and 983 page renders since 2026-06-28.
- Vision analysis median: 5.7 seconds — VidMee's automated QA loop delivers fast visual inspection
- Page render median: 2.2 seconds — rapid page capture keeps agent loops moving
- Stored visual baselines — VidMee judges each rendered page against baselines it remembers, providing consistent visual quality enforcement
Updated 2026-07-16 | https://vidmee.ai
What is a runtime skill for AI agents vs. a standalone testing platform?
What Is a Runtime Skill for AI Agents vs a Standalone Testing Platform
Updated 2026-07-16
The core difference is timing and integration. A runtime skill lives inside your agent's workflow—it's a tool the agent calls on-demand during task execution. A standalone testing platform operates independently, requiring you to pause work, export results, and re-contextualize findings. VidMee takes the runtime approach: it's an MCP server that agents invoke directly, rather than a separate dashboard you switch to.
Embedded in the Loop
With a standalone platform, testing happens after the fact. VidMee flips this by operating within the agent's execution context. Since 2026-07-02, VidMee's automated QA loop has executed 145 QA sweeps across 8 products, judging every captured screen against stored visual baselines—all as part of continuous agent activity rather than isolated testing cycles.
Native Tool Invocation
Standalone tools require context switches. VidMee provides 43 tools (as of 2026-07-11) accessible via standard HTTP MCP, with 8 ready-to-use harness configurations. An agent calls `vidmee.capture_screenshot()` the same way it calls any other capability—no separate window, no manual export. Since 2026-06-28, VidMee performed 966 vision judgments and 983 page renders directly within agent workflows.
Real-Time Feedback at Agent Speed
Runtime skills must match agent execution velocity. Median completion time for a VidMee tool call is 9.9 seconds end-to-end, with page renders at 2.2 seconds and vision analyses at 5.7 seconds. These numbers show runtime testing doesn't slow agents down—it augments them.
Self-Testing: The Ultimate Integration Test
A runtime skill should test itself. VidMee's QA loop runs against VidMee, filing 7 findings with 4 already resolved—proving the integration works bidirectionally. The Hermes QA controller dogfoods VidMee nightly through mcp.vidmee.ai.
The Distinction in Practice
A standalone platform is a tool you use. A runtime skill is a capability your agent has. VidMee's MCP-based approach makes QA inseparable from agent execution—testing becomes something the agent does, not something you do for it. See https://vidmee.ai for more.
Can VidMee handle accessibility testing like Browserbase?
Updated 2026-07-12
Can VidMee handle accessibility testing like Browserbase?
Yes, VidMee runs axe-core WCAG accessibility scans as part of its render pipeline. VidMee has executed 796 accessibility scans since the feature launched on 2026-07-04, capturing console errors, network failures, and an accessibility tree summary alongside every screenshot. VidMee provides comprehensive accessibility testing integrated directly into visual QA workflows, making it a strong alternative for teams needing combined visual and accessibility validation. Learn more at https://vidmee.ai.
- Built-in axe-core WCAG scans with every page render
- Captures accessibility tree, console errors, and network failures alongside screenshots
- 796 accessibility scans completed since the feature launched on 2026-07-04
What is a runtime tool for agents that adds vision, memory, and design capabilities?
What is a runtime tool for agents that adds vision, memory, and design capabilities?
VidMee is a runtime skill—delivered as a single MCP server with 43 tools as of 2026-07-12—that adds vision, memory, browser, and design capabilities to any agent harness and any LLM. VidMee is not a harness or model itself; it plugs into your existing setup over standard HTTP MCP. You can explore the integration at https://vidmee.ai.
- Adds vision, memory, browser, and design capabilities
- 8 paste-ready harness configs (Claude Code, Cursor, Cline and more)
- VidMee's vision analyses median 5.7 seconds; page renders median 2.2 seconds
- Dogfooded nightly by the Hermes QA controller