Tag: MCP

  • Model Hardware Standard: Anthropic Cuts Lab Setup to 8 Hours

    Anthropic released the Model Hardware Standard on August 27, 2026 — a research preview that lets Claude and rival models drive lab robots, pipettes and factory arms through a single spec. Early testers cut integration from weeks to hours. Carnegie Mellon stood up a serial dilution workflow in 8 hours. QuEra took a laser recovery routine from 58% success to 99.3%. No pricing, no revenue, no open-source date.

    Anthropic has spent two years selling tokens that move text. This one moves matter.

    The company published the Model Hardware Standard, or MHS, as a research preview on August 27. It is a specification, not a product — closer to a plug shape than to a machine. And that is exactly the point.

    What is the Model Hardware Standard?

    The Model Hardware Standard is a shared specification that tells an AI agent what a physical device can do and, more importantly, what it must never do. Vendors ship a driver. The agent reads and writes through simple primitives. Anthropic is running it as an invitation-only research preview.

    Per Anthropic’s own announcement, MHS works with any device that exposes a programmable interface. It is model-agnostic by design — Claude is not required.

    That last detail matters more than the demos. Anthropic is not shipping a robot. It is trying to own the socket every robot plugs into.

    How MHS actually works

    A vendor writes one standardized driver. That driver publishes device discovery in a common format and exposes controls, sensor values and safety limits through a shared memory dictionary.

    The agent then reaches the hardware through one of three paths: the Model Context Protocol, a command line interface, or generated code files. Same device, three levels of abstraction.

    Safety limits live in the driver, not in the prompt. Anthropic gives the example of blocking excess laser power at the device layer — so a confused model cannot talk its way past a hardware ceiling.

    Where MCP ends and the Model Hardware Standard begins

    MCP, which Anthropic debuted in 2024, connects models to software: databases, ticket systems, file stores. If you have followed our coverage of how Agent Skills and MCP split the token bill, the architecture will look familiar.

    MHS extends the same logic to things with motors. Anthropic technical staff member Alek Kemeny put it bluntly to TNW: “What MCP did for software, MHS will do for the hardware world.”

    Kemeny has described MCP elsewhere as “kind of like the USB for AI to software connection.” MHS is the industrial-grade version of that pitch.

    What did the early tests actually prove?

    Six organizations ran MHS against real equipment before launch, and the reported results are specific rather than vague. The headline claim is time: Anthropic says MHS “reduces this integration work to hours or minutes,” against a baseline Genentech described as weeks or months of manual work.

    The most concrete number came from quantum computing. QuEra used MHS to rebuild a laser stabilization routine, moving from 58% success at 150 seconds to 99.3% success at 6 seconds — a 25x speedup on a task that was already mostly failing.

    Organization What was automated Reported result
    QuEra Computing Laser stabilization recovery 58% to 99.3% success; 150s to 6s
    Carnegie Mellon Serial dilution workflow 3x faster; 8-hour setup vs. weeks
    Tetsuwan Scientific qPCR liquid handling 9,143 dispenses across 300 transfer types
    Genentech BCA protein assay tuning Converged at ~140 µL/s (water), 10 µL/s (BSA)
    University of Washington Multi-instrument bench Six instruments connected in under a week
    HHMI Janelia Co-development partner Reference implementation

    Anthropic also says it tested six failure conditions on purpose: missing plate, rotated plate, reader busy, disconnected camera, unreachable device, emergency stop. That is a short list for anything touching a factory floor.

    The number that should give buyers pause

    Every figure above comes from partners Anthropic selected and published. None of it is independently benchmarked, and there is no public failure rate across the full preview cohort.

    A 99.3% success rate on a laser is excellent in a lab. On a production line running 20,000 cycles a shift, it is 140 faults.

    Who is backing the Model Hardware Standard?

    Anthropic named ten hardware vendors and six research institutions at launch. The vendor list is the commercially interesting half, because those are the companies that would have to ship MHS drivers in firmware for the standard to matter.

    Vendors listed by Anthropic as supporting or planning support:

    • Amazon Web Services (Strands Robots library)
    • Universal Robots
    • Doosan Robotics
    • Danaher
    • QIAGEN
    • Tecan
    • Automata
    • MBF Bioscience
    • Hugging Face (LeRobot)
    • Raspberry Pi

    Research users include Genentech, Carnegie Mellon, the University of Washington’s Baker and Pinglay labs, HHMI Janelia, QuEra and Tetsuwan Scientific.

    Jonah Cool, Anthropic’s head of partnerships and deployment of science, told Fortune that lab equipment “suffers from proprietary solutions that are very brittle,” and that the goal is to “avoid vendor lock-in for scientists.”

    Read that again from a vendor’s chair. Anthropic is asking Danaher, QIAGEN and Tecan to help dismantle the integration moat that protects their service revenue.

    How much does the Model Hardware Standard cost?

    Nothing, for now — and that is the strategy. MHS is free during the research preview, gated by an invitation waitlist at modelhardwarestandard.com. Anthropic says it will open-source the framework after the preview, but has published no date, no license and no commercial terms.

    Standards are loss leaders. The money is downstream, in the tokens burned by agents that run instruments around the clock.

    An overnight experiment is a 12-hour inference session. Multiply that by a few thousand labs and the economics start to look like a metered utility rather than a chat subscription.

    Who wins and who loses financially?

    The winners are frontier labs with agent products and the robotics vendors with thin software teams. The losers are instrument makers whose margins depend on proprietary integration, and the systems integrators paid by the week to wire benches together.

    Winners

    Anthropic first. The company was reported at a $2 trillion valuation earlier this month, and a hardware standard extends its distribution into a market where it currently sells nothing.

    Robot arm vendors win cheaply. Universal Robots and Doosan get an agent interface without building an AI stack — the same trade that made Unitree’s IPO pop 629% a bet on hardware plus somebody else’s brains.

    Cloud providers win the runtime. AWS shipped Strands Robots support on day one for a reason.

    Losers

    Integration consultancies are the clearest casualty. If a Carnegie Mellon bench goes from several weeks to 8 hours, that is billable work evaporating.

    Proprietary lab software is next. MarketsandMarkets valued lab automation at $6.60 billion in 2026, growing to $8.62 billion by 2031 at a 6.6% CAGR — a slow market where vendors defend share through lock-in, not growth.

    A commoditized driver layer is precisely the thing that breaks that defense.

    Is the Model Hardware Standard safe enough to run a factory?

    Not yet, and Anthropic says so. The company acknowledged that large language models “still lack physical intuition,” and states that safety evaluations are being built during the preview rather than before it. Human approval workflows exist for high-risk actions, but the physical safety roadmap is unfinished.

    The Register, which covered the launch on August 28, raised the obvious dual-use question: a universal spec for driving instruments does not care what the instrument is for.

    The January 2027 regulatory deadline

    EU Machinery Regulation 2023/1230 takes effect on January 20, 2027. It is the first EU rule to cover AI-based safety functions and self-evolving machine behavior.

    TNW notes the awkward implication: an MHS file that constrains how a machine may operate could itself qualify as a regulated safety component. That would put liability on whoever wrote the driver.

    Anthropic has not said who that is. Five months out from the deadline, this is the unpriced risk in the whole announcement.

    How does this fit Anthropic’s broader agent push?

    MHS is the physical endpoint of a strategy that has been visible all year in software. Anthropic has been widening what an agent can touch, from computer-use agents driving desktops to skills that compress tool definitions.

    The competitive timing is not subtle either. Fortune reported that Hugging Face shipped a robotic duck the same day, and that Nvidia is pursuing a $13 billion acquisition of the company.

    Physical AI is where the capital is rotating. German humanoid maker NEURA Robotics raised up to $1.4 billion in Series C funding this year, per TNW.

    Frequently asked questions

    Is the Model Hardware Standard open source?

    Not yet. Anthropic says it intends to open-source the framework after the research preview, but has published no date or license. Drivers built during the preview are being made available for reuse.

    Does MHS only work with Claude?

    No. Anthropic describes MHS as model-agnostic, meaning OpenAI models and open-weight models can drive MHS devices. Whether rival labs adopt a spec authored by a competitor is a separate question.

    How is MHS different from MCP?

    MCP connects models to software. MHS connects them to physical devices, and adds device-level safety limits, sensor state and discovery. MCP is one of three ways to reach an MHS device, alongside a CLI and generated code.

    Can I use it today?

    Only by invitation. Access runs through a waitlist at modelhardwarestandard.com, and Anthropic has described early access as a “handful” of labs and manufacturers in biotech, robotics and quantum computing.

    What hardware is supported?

    Anything with a programmable interface, in principle. In practice, ten named vendors — including Universal Robots, Danaher, QIAGEN, Tecan and Raspberry Pi — are supporting or planning support. Older instruments without a programmable interface are out of scope.

    What is the biggest risk?

    Regulation and liability. EU Machinery Regulation 2023/1230 applies from January 20, 2027, and MHS constraint files may count as regulated safety components — with no clarity yet on who carries responsibility when an agent-driven machine injures someone.

    The bottom line

    The Model Hardware Standard is the most strategically aggressive thing Anthropic has shipped this year, and it contains no product.

    The engineering claims are credible and unusually specific. A 58% to 99.3% jump on QuEra’s laser routine and an 8-hour Carnegie Mellon integration are not marketing numbers. They are the kind of figures a skeptical buyer can go test.

    But every one of them came from a partner Anthropic chose. There is no pricing, no open-source date, no independent benchmark and no answer on who is liable when a driver written by a language model moves a robot arm into a person.

    The verdict: treat MHS as a distribution land-grab, not a revenue event. If ten vendors becomes fifty by January, Anthropic will own the plug shape for physical AI and collect inference rent on every machine that uses it. If the EU deadline arrives with the liability question still open, the same vendors will quietly wait it out.

    Watch the driver count, not the demos.

    Sources

  • Agent Skills vs MCP: Which One Cuts Your Token Bill

    Agent Skills vs MCP is not really a fight — but on cost, Skills win outright. Anthropic’s own numbers show 58 MCP tools burning roughly 55,000 tokens before an agent does anything useful. A Skill’s metadata costs about 100 tokens. Build Skills for procedure and judgment, MCP servers for live system access, and cache aggressively if you run both.

    On August 19, 2026, Anthropic moved Agent Skills, the Skills API, computer use, browser use and the Files API out of beta and into general availability on the Claude Developer Platform, according to the platform’s release notes. Three weeks earlier, the Model Context Protocol shipped its stateless 2026-07-28 specification.

    Both are now production infrastructure. Both are free to adopt. Only one of them charges you rent on every single request.

    What actually changed on August 19, 2026?

    Agent Skills stopped being an experiment. The skills-2025-10-02 beta header is gone, the /v1/skills endpoint is GA, and computer use shipped as computer_toolset_20260801 with browser use as a separate tool, browser_toolset_20260801. That is a full production agent stack in one release.

    The Files API went GA the same day. Upload, download, list, metadata and delete operations are free — you are billed only for file content that actually enters a Messages request, at standard input-token rates.

    Storage caps are generous: 500 MB per file and 1 TB per organization, rate-limited to roughly 500 requests per minute.

    None of the superseded beta headers have a published sunset date yet. Legacy identifiers still work. But GA is the signal that pricing and architecture decisions made today will stick.

    What is the difference between Agent Skills and MCP?

    MCP gives an agent reach. Agent Skills give an agent method. An MCP server tells Claude how to connect to GitHub and what it can do there. A Skill tells Claude how your team actually writes a release note, in what order, with which checks.

    That distinction sounds academic until you look at where each one lives in the context window.

    How MCP loads tools

    MCP tool definitions are pushed into the model’s context up front, on every request. The agent has to know a tool exists before it can call it, so the full schema — names, parameters, descriptions — sits in the prompt whether or not the tool ever gets used.

    That is a fixed tax. It scales linearly with how many servers you connect.

    How Agent Skills load

    Skills use three-tier progressive disclosure, documented in Anthropic’s Agent Skills overview. Level 1 is YAML frontmatter — name and description only — at roughly 100 tokens per Skill, always loaded. Level 2 is the SKILL.md body, under about 5,000 tokens, loaded only when the Skill is triggered.

    Level 3 is where it gets interesting. Bundled reference files and scripts cost zero tokens until read. Script code never enters the context window at all — only its output does.

    You can ship 300 pages of API documentation inside a Skill and pay nothing for it unless the agent opens the file.

    How much do MCP tool definitions actually cost?

    More than most teams realize. Anthropic published a concrete five-server example in its advanced tool use write-up: GitHub at 35 tools (~26K tokens), Slack at 11 tools (~21K), Sentry at 5 (~3K), Grafana at 5 (~3K), and Splunk at 2 (~2K).

    That is 58 tools consuming approximately 55,000 tokens before the conversation even starts.

    Anthropic’s own internal setup is worse: tool definitions there consume 134,000 tokens before optimization. On a 200K context window, that is two-thirds of the budget spent on a menu the agent mostly ignores.

    Now price it. Claude Sonnet 5 costs $2 per million input tokens and Claude Opus 5 costs $5, per the official pricing page. Sonnet 5’s introductory rate was made permanent on August 10, 2026, and the scheduled September increase was cancelled.

    What does that cost per month at real volume?

    Setup Tokens per request Sonnet 5 cost / request 10,000 runs / month
    Anthropic’s internal tool set 134,000 $0.268 $2,680
    58 MCP tools (5 servers) ~55,000 $0.110 $1,100
    Same tools + Tool Search Tool ~8,700 $0.017 $174
    20 Skills + 1 triggered SKILL.md ~7,000 $0.014 $140
    20 Agent Skills (metadata only) ~2,000 $0.004 $40
    Token figures from Anthropic; cost math at the published $2/MTok Sonnet 5 input rate, uncached.

    The spread between the top and bottom row is 27x. On Opus 5 at $5 per million input tokens, the same gap costs $6,700 versus $100 a month.

    The prompt caching escape hatch

    MCP defenders have a real counterargument: cache the tool definitions. Anthropic’s pricing page lists cache reads at 0.1x the base input rate — $0.20 per million tokens on Sonnet 5.

    Cached, that 55,000-token block drops from $0.110 to about $0.011 per request. A 1-hour cache write costs 2x base and pays for itself after two reads.

    So caching closes most of the gap — if your traffic is dense enough to keep the cache warm and your tool list is stable. Bursty, low-volume agents get cache misses and pay full freight.

    Does trimming tools hurt accuracy?

    No — it helps. This is the part that surprises people. Anthropic’s Tool Search Tool delivers an 85% reduction in token usage while keeping the full tool library reachable, cutting roughly 77K tokens of overhead down to about 8.7K and preserving 95% of the context window.

    Accuracy went up. Opus 4 improved from 49% to 74% on the tool-use benchmark. Opus 4.5 went from 79.5% to 88.1%.

    Programmatic Tool Calling shows the same pattern: average usage dropped from 43,588 to 27,297 tokens, a 37% reduction on complex research tasks, while internal knowledge retrieval rose from 25.6% to 28.5% and GIA scores went from 46.5% to 51.2%.

    Fewer tools in context means less for the model to confuse. Context bloat is an accuracy problem wearing a cost problem’s clothes.

    When should you still build an MCP server?

    When you need a live connection, real authentication, or one integration shared across many agents. Skills are static files — they cannot hold an OAuth token, stream an update, or talk to your database. MCP is the transport layer, and nothing about Skills replaces it.

    MCP’s adoption numbers back that up. Claude’s connector directory now lists over 950 MCP servers, and MCP has passed 400 million monthly SDK downloads — a 4x increase this year, per Anthropic’s spec announcement. The TypeScript and Python SDKs have crossed a billion total downloads between them.

    Use this five-question test before you write a line of either:

    1. Does it need a live connection to a running system? MCP server.
    2. Does it need per-user auth or scoped permissions? MCP server.
    3. Is the hard part procedure and judgment, not access? Agent Skill.
    4. Is it large reference material used occasionally? Agent Skill — Level 3 files cost nothing until opened.
    5. Do many agents share one integration? MCP server centralizes it; a Skill travels with the agent.

    Agent Skills vs MCP: which should you use for your task?

    What you’re building Pick Reason
    Read/write live Slack, GitHub or Jira data MCP server Needs a live, authenticated connection
    House style guide, review checklist, report format Agent Skill Pure procedure; ~100 tokens idle
    Ship 300 pages of API reference to the agent Agent Skill Level 3 files cost 0 tokens until read
    Per-user OAuth scopes and audit trails MCP server Auth belongs at the connection layer
    Deterministic script the agent runs but shouldn’t read Agent Skill Script code never enters context
    One integration consumed by a dozen agents MCP server Central updates, single surface
    Pull live data and apply a fixed workflow Both MCP for access, Skill for method
    Low-volume, bursty agent on a tight budget Agent Skill Cold caches make MCP overhead expensive
    The production default is both — MCP for reach, Skills for method.

    What did the 2026-07-28 MCP spec change?

    It made MCP stateless, which is the single biggest cost change on the server side. The initialize/initialized handshake and the Mcp-Session-Id header are retired. Every request is now self-describing, so any request can land on any instance behind a plain round-robin load balancer.

    That kills the shared-storage requirement. You can run MCP serverless and stop paying for sticky sessions.

    Method and tool names now travel in HTTP headers rather than JSON bodies, letting gateways route without parsing payloads. Multi Round-Trip Requests replace server-initiated streams for interactive confirmations.

    Roots, Sampling and Logging are deprecated but will keep working for at least twelve months, as will the legacy HTTP+SSE transport. You have a year to migrate — not a weekend.

    One number worth sitting with: at Honeycomb, nearly 20% of all monthly interactive queries are now made by agents, not humans. That ratio is why the per-request tax matters.

    Frequently asked questions

    Are Agent Skills a replacement for MCP?

    No. Skills carry procedural knowledge as files; MCP carries live, authenticated access. Anthropic shipped both to GA in 2026 and the standard production pattern uses them together.

    How many tokens does one Agent Skill cost?

    Roughly 100 tokens for its metadata at startup, per Anthropic’s documentation. The SKILL.md body — under about 5,000 tokens — loads only when the Skill is triggered.

    Do MCP tool definitions get charged on every request?

    Yes, unless cached. Anthropic’s example of 58 tools across five servers consumes roughly 55,000 input tokens per request. Prompt caching cuts that to 0.1x the base rate on cache hits.

    Does the Files API cost extra?

    No. Upload, download, list, metadata and delete are free. You pay only for file content that enters a Messages request, billed as normal input tokens.

    Is the old MCP spec still supported?

    Yes. Roots, Sampling, Logging and the HTTP+SSE transport are deprecated but supported for at least twelve months from the 2026-07-28 release.

    Which model should I run agents on to keep costs down?

    Claude Sonnet 5 at $2/$10 per million tokens is the volume workhorse; its introductory pricing was made permanent on August 10, 2026. Opus 5 at $5/$25 is 2.5x the input cost — worth it only when the task genuinely needs it.

    Do fewer tools in context make agents dumber?

    The opposite. With Tool Search Tool enabled, Opus 4.5 improved from 79.5% to 88.1% on Anthropic’s tool-use benchmark while using 85% fewer tokens.

    The bottom line

    Build Skills first. They are nearly free to keep loaded, they went GA on August 19, and Anthropic’s own benchmarks show that trimming context raises accuracy rather than lowering it.

    Add MCP servers only where you need a live connection, real auth, or one integration shared across many agents — then wrap them in Tool Search Tool and prompt caching on day one, not after the first surprising invoice.

    The decision rule is exactly this: if the hard part is reaching the system, build MCP; if the hard part is knowing what to do once you’re there, build a Skill. Teams running high-volume agents on the naive pattern are paying somewhere between 8x and 27x more per request than they need to, and getting worse answers for the money.

    Related reading on agent economics: our breakdown of Claude Code vs Codex CLI on cost, the Gemini 3.7 Flash vs Claude Sonnet 5 cost-per-coding-point comparison, and ChatGPT Business vs Claude Team on the seat-license side.

    Sources