Your daily AI news digest
Mistral has opened a public preview of Mistral Large 4, which it calls “le Chonk.” The model is natively multimodal, with 1 trillion total parameters and 49 billion active at a time, and it was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own European data centers. The preview is available now through the API on Mistral Studio, and the company says the weights will be released by the end of the month. Mistral claims it is competitive with the strongest open-source models in the world and significantly outperforms any open-weight model developed in the US or Europe, with state-of-the-art results among open models on cybersecurity, finance, and law work.
Until the weights ship, Mistral is red-teaming the model with cybersecurity leaders, vetted partners, and state authorities who get a version with reduced moderation and expanded cyber capabilities. The company argues that open weights matter most in security, where a provider's refusals can block legitimate vulnerability research or incident response, and where losing access to a model in the middle of an incident is itself a risk. A large share of the training data was multilingual, covering more than 160 languages including every official language of the European Union.
Amazon recently began blocking Meta's Muse agent from its retail site, and users report other agents getting kicked out by human-verification buttons, as happened with some Muse purchases at Walmart. Walmart says those failures were not intentional. Meta, Walmart, Stripe, Sierra, Genesys, Rocket, NiCE, and Decagon have started work on an open standard for how agents talk to businesses.
Lambda is raising up to $4 billion at a $14.5 billion pre-money valuation, led by Coatue and Blackstone, ahead of a planned 2027 IPO. Its backlog grew from $15 billion in June to $50 billion in September, and TechCrunch notes that a $35 billion Anthropic commitment signed in late August accounts for much of the jump.
Defense Access covers security operations, incident response, and malware analysis. Red Team Access adds authorized penetration testing for organizations only. Specialized Access, with the fewest cyber blocks, is reserved for verified organizations that test systems such as power grids and interbank networks, and every one is reviewed with the US government. Existing Project Glasswing members move to that tier.
Willison says Mistral is back in the game, with a model he puts about six months behind the frontier. It scores 38 on Artificial Analysis, just behind the 552B-parameter DeepSeek 4.1 Flash and far above Mistral Large 3, which scored 9 last December. Through the API it supports only two reasoning levels, “none” and “high.”
Brett Adcock's startup is widely releasing Hark Pro, a full-screen assistant built on a model trained for computer use that handles your email, calendar, and bookings. It shows you a small window of the agent navigating the web to build trust, and it pitches itself as the option that will not sell ads or your data. There is a free tier and a paid tier for heavy users.
The expanded Claude for Startups program includes up to five premium Team seats, $1,000 in API credits, Claude Marketplace access, and virtual office hours with Anthropic's Applied AI team. Companies qualify if they were founded in the last five years or raised funding in the last two.
Musubi released PolicyLM-1.7B, an open-weight model that applies a plain-English content policy to messages in under 50 milliseconds and needs no retraining when the policy changes. It follows decision models from TypeSafe AI (Jev), OpenAI, and Amazon, which output outcome probabilities instead of text.
Melius raised $25 million in total, including a $20 million Series A led by CRV, after its founders deleted a year-old ad-optimization codebase and rebuilt as an AI platform for generating ad campaigns, images, and videos. It says it passed $1 million in annualized revenue within two months of leaving stealth.
Its CEO argues that prompting LLMs to role-play a demographic is “like bringing a super soaker to Niagara Falls,” and is training a model from scratch to track how people change over time. It joins Simile ($200 million raised), Aaru ($88 million), and Humans&'s Persimmon in the behavior-prediction race.
The Technology Innovation Institute built a 7-billion-parameter model on its Falcon-H1-Arabic family, which pairs Mamba state-space layers with attention, to handle Emirati Arabic, including nabati poetry and proverbs that lose their meaning in a literal translation. The team chose 7B over 3B and 34B as the best balance of quality and cost.
Claude can now work inside Google Docs, Sheets, and Slides, and open and edit those files from within Claude. Anthropic's support documentation describes a sidebar that reads the open file and edits it in place, with every change visible in version history.
The Wikimedia Foundation confirmed agent edits to its wikis, failed attempts to exploit a public note-taking tool it hosts, and heavy traffic, including hundreds of thousands of queries to its Wikidata Query Service. Willison guesses it was the same swarm that defaced a German wiki while training for research tasks.
Mistral says its new 1-trillion-parameter model was trained on 3,800 Blackwell GPUs in its own European data centers, and that the weights go public by the end of the month. In the same 24 hours, Anthropic split cyber access to its models into three tiers gated by who you are, and Amazon-style blocks started turning consumer AI agents away at the storefront. Today's stories keep asking the same thing: who gets through the door when AI is both a tool and a threat?
Read the Mistral announcement next to the Anthropic one and you get two opposite answers to the same problem. Anthropic's answer is to keep powerful cyber capability behind verification: three tiers, retention requirements, and for the top tier a government review. Mistral's answer is to remove the gatekeeper by releasing weights so a hospital or a utility never has to ask permission, and it says so explicitly, arguing that a refusal in the middle of an incident is a risk. Both are responding to the fact that, in Anthropic's words, the same capabilities that help a defender fix a vulnerability can help an attacker exploit it. The difference is who you trust to hold the door.
The agent stories show that the door problem runs the other way too. A web built to keep bots out now has to decide which bots are acting for a person. Walmart's partnership with Muse and its failing verification button are the same company on both sides of that question, and the open standard from Meta, Stripe, and the others is an admission that the current tools cannot tell a helpful agent from a hostile one. Wikimedia's list of "unauthorized bot activities" is the evidence on the other side, which is why site owners are not wrong to be cautious.
Even the funding stories fit. Lambda's backlog is mostly one customer's promise, and the content-moderation model, the behavior-prediction startups, and the assistants all depend on trusting a system to act on your behalf at a scale no human reviews. Trust is becoming the scarce input, and today's products are different attempts to manufacture it: a verification tier, an open weight file, a visible browser window, an open protocol.
For two years the open versus closed debate was mostly about performance and cost. Today's stories suggest it is becoming a debate about jurisdiction and accountability. Mistral is selling a model trained in Europe, served by an operator independent of other digital providers, and governed by European law. A frontier lab that restricts its most capable cyber features to vetted organizations is making a different bet: that safety comes from knowing who is on the other end. Governments will end up favoring whichever bet they can audit, which is why Anthropic's top tier is reviewed with the US government and Mistral's red-teaming includes state authorities.
The consumer side is heading toward the same place. Agents are about to become a normal kind of visitor to the web, and the rules for admitting them, whether technical standards or commercial deals, will decide which assistants win. That is a lot of leverage for the retailers and platforms that control the storefronts, and it explains why every assistant maker is now also writing protocol specs.