Brett Adcock's AI Lab Launches Handoff: A Browser Agent Built for Everyday Tasks, Not Code
Hark Labs, the AI startup founded by serial entrepreneur Brett Adcock, announced Handoff on August 5, 2026, a browser-use agent designed to autonomously handle everyday web tasks like ordering food, booking flights, and shopping. The model scored 97.7 on the Online-Mind2Web benchmark, a third-party evaluation for web agents, compared to 92.8 for OpenAI's GPT 5.4 and 84.1 for Anthropic's Claude Opus 4.8. Handoff will enter research preview at hark.com with wider availability planned for later this summer.
What Makes Handoff Different From Other AI Agents?
While most AI agents in development focus on coding and developer workflows, Handoff explicitly targets consumer and business tasks. Adcock stated: "While others focus on coding, we focus on everyday life: ordering food, booking flights, shopping, and navigating the web". This positioning distinguishes Handoff from coding-focused agents like OpenAI's Codex or Anthropic's Claude Cowork, which are built for developers working in code editors and terminals.
Adcock
The practical motivation behind this approach is straightforward. Hark's research found that despite people spending 75 percent of their screen time daily in a browser, fewer than 1 in 1,000 websites have publicly accessible APIs. This gap means AI agents cannot rely on direct integrations to complete web tasks; they must instead learn to navigate websites the way humans do, clicking buttons, filling forms, and reading text on screen.
How Does Handoff's Pricing Compare to Competitors?
One of Handoff's clearest advantages is cost. Hark says the model can be served at less than one-tenth the token price of competing frontier models, charging $0.18 per million input tokens and $2.37 per million output tokens. By comparison, OpenAI's GPT 5.5 costs $5 per million input tokens and $30 per million output tokens, and Anthropic's Opus 5 carries the same pricing as its predecessor. Handoff also claims per-turn latency of 0.8 seconds, significantly faster than the 6 to 6.8 seconds Hark measured for competing systems.
However, important caveats apply to these performance claims. The latency comparisons were measured by Hark using Hark's own testing setup, with competing models set to their highest reasoning level, which is their slowest configuration. No independent third party has verified these latency measurements.
What Are the Limitations of Handoff's Benchmark Claims?
Hark's benchmark comparisons focus on prior-generation models rather than the current frontier. The company compared Handoff against GPT 5.4, GPT 5.5, Opus 4.8, and Gemini 2.5 Pro, but notably absent are OpenAI's GPT 5.6 and Anthropic's Opus 5, the current leading models. These newer systems have not yet published results on the Online-Mind2Web benchmark, so Hark's "top-ever" claim cannot be verified against the strongest available systems.
This omission matters because the newest frontier models have posted their largest gains precisely in computer use tasks. On OSWorld 2.0, another related benchmark covering full computer control, Anthropic's Opus 5 scores approximately 70.6 percent compared to 55.7 percent for the Opus 4.8 model Hark selected for comparison. Additionally, on WebTailBench v2, one of Hark's own three chosen benchmarks, GPT 5.5 actually scores 72.3 compared to Handoff's 68.6.
Two of the three benchmarks Hark cited were also run inside Hark's own testing environment, with pass rates computed by Hark's internal language model judge, meaning the company controlled the evaluation conditions. When asked whether Hark plans to publish comparisons against newer models, the company did not specify.
How Was Handoff Built, and What Remains Unknown?
Hark's research preview describes a reasonable training pipeline: supervised fine-tuning followed by asynchronous reinforcement learning using the GRPO algorithm. However, the company acknowledges it has only completed post-training so far, with pre-training planned for later this year. This means Handoff is built on top of a base model that Hark did not train itself. When asked which base model it uses and what mix of proprietary and open data Handoff was trained on, Hark has not yet specified.
For enterprise users, significant questions remain about security and data access. A Hark spokesperson said "security and privacy is a primary focus, but this is a technical preview," adding the company will share more details when the product reaches market at the end of summer. Specifically, it remains unclear who can access the dedicated virtual computers and files created on them when users run tasks through Handoff.
Steps to Understanding Handoff's Place in the AI Agent Landscape
- Benchmark Context: Handoff's top scores come from comparisons against prior-generation models, not current frontier systems like GPT 5.6 or Opus 5, making independent verification against the strongest competitors impossible at this time.
- Cost Advantage: The model's pricing of roughly one-tenth that of competing systems is Hark's clearest competitive claim and would hold even against current frontier models if performance claims are verified.
- Technical Transparency: Key details remain undisclosed, including the base model used, the data mix for training, and security protocols for enterprise deployments, all of which will be critical for adoption.
- Product Maturity: Handoff is in research preview with wider availability promised later this summer, meaning real-world performance data from independent users is not yet available.
What Does This Mean for Brett Adcock's Broader Ambitions?
Hark represents Adcock's fourth company. He previously co-founded Vettery, a talent marketplace sold in 2018 for roughly $100 million, and Archer Aviation, an air-taxi maker. He also founded Figure AI, the humanoid robotics company that has become a unicorn. Adcock raised a $700 million Series A round for Hark in May 2026 at a $6 billion valuation, led by Parkway Venture Capital with participation from Nvidia, AMD, Intel Capital, Qualcomm Ventures, Salesforce Ventures, and ARK Invest. Adcock seeded the company with $100 million of his own money and remains founder and CEO of both Figure and Hark simultaneously.
When asked how the two companies interact, a Hark spokesperson said Hark models "are being trained on the Figure robots," but that Adcock has no plans to combine them. This suggests Hark may be developing AI systems that benefit from robotics data, though the exact nature of that relationship remains unclear.
Adcock's promotional style has drawn scrutiny in the past. In April 2025, Fortune correspondent Jason Del Rey reported that Figure's much-touted BMW partnership was far more modest than Adcock's public claims suggested, with BMW confirming only a single robot was practicing parts pickup during non-production hours. However, the partnership has advanced; as of June 2026, BMW said the Figure 02 robot supported production of more than 30,000 BMW X3 vehicles over a 10-month period, and the next-generation Figure 03 robot was being deployed for parts-sequencing in logistics. On social media, Adcock called the earlier Fortune story "mischaracterizations and downright lies" and threatened a defamation suit.
None of this necessarily means Handoff's numbers are wrong. The agent may well perform excellently, and the pricing, if it holds, would undercut every major lab. But the pattern suggests caution is warranted until independent reviewers can test Handoff themselves during the research preview phase.