Moonshot announced what it calls the first open model of its size, and did not release the weights.

Kimi K3 went live on July 16: a 2.8-trillion-parameter sparse mixture-of-experts model activating 16 of 896 experts per token, a 1-million-token context window, image input, and a new attention stack. It was available to use that day on kimi.com and the Kimi API. The announcement is titled Open Frontier Intelligence.

Read Moonshot’s own sentence, though: “The full model weights will be released by July 27, 2026.” That is a promise with a date on it, and the date is after this week ended. No license has been named for the eventual release. No training code or data is promised.

So for the week it was the most discussed model release on Hacker News, “open” meant a promissory note with a date on it.

The day before, a different lab did the other thing.

2,082 / 1,218

The announcement got 2,082 points. The actual weights got 1,218.

Kimi K3, announced without weights, took 2,082 points on Hacker News. Inkling, Thinking Machines’ 975-billion-parameter multimodal model released under Apache 2.0 with the files downloadable the same day, took 1,218.

One of those you can download and run yourself. The other was available only as hosted inference.

(Live counts as of Monday morning, July 20, shortly before this issue went out. These are moving counters and a snapshot, not a week-long measurement, and points count votes on a submission rather than readers.)


🧩 What “Open” Bought You This Week

Two releases on consecutive days, and only one of them is a file you can hold

Inkling is 975B total parameters with 41B active, pretrained from scratch on text, images, audio and video, under a named OSI license with day-zero support in vLLM, SGLang and Unsloth. The artifact those claims attach to is on disk right now.

And here is the number that keeps this honest, because it is not the flattering one. Artificial Analysis scored Inkling at 41 on the same index where it scored K3 at 57. Inkling is the leading open-weights release from a US lab, three points ahead of Nemotron 3 Ultra, and it is sixteen points behind the model that did not ship. So this is a real tradeoff, not a morality tale. The weights you can download are meaningfully weaker than the weights you cannot.

K3 scores higher on the one independent evaluation available. Artificial Analysis, an evaluator independent of Moonshot, scored it 57 on Intelligence Index v4.1. Configurations matter on that leaderboard and we should name them: it sits behind Claude Fable 5 (with fallback) at 59.9 and GPT-5.6 Sol (max) at 58.9, and ahead of GLM-5.2 (max) at 51.1 and DeepSeek v4 Pro at 44. That Fable 5 parenthetical is worth stopping on, because it is not an effort tier like the others. Per Anthropic’s own cookbook, Fable 5 ships with safety classifiers that run on every request and block anything touching offensive cybersecurity or the life sciences. Anthropic’s guidance is that API customers configure a fallback to Claude Opus 4.8 for those cases, via a server-side beta or client-side SDK logic. Configured, the blocked request lands on Opus 4.8; unconfigured, it comes back as a refusal. Anthropic’s words, not ours: the safeguards “limit its performance in these specific areas,” they are “deliberately conservative,” and “benign technical work sometimes triggers them.”

So the top row of that leaderboard is, by configuration, two models. How much it mattered to the score is a fair question we cannot answer: if the eval set is light on cyber and biology, the fallback may have rarely fired. The configuration permits the handoff; the extent is unmeasured. Credit to Artificial Analysis for labelling it at all, which is the only reason anyone can ask the question. That is a genuinely strong result for a model you could not download while it was being tested, and Artificial Analysis says so explicitly in its own writeup.

Two things in that same evaluation did not make the headlines.

The price went up, a lot. K2.6 cost $0.95 per million input tokens and $4.00 output. K3 costs $3.00 and $15.00. That is 3.1x input and 3.75x output, generation over generation. K3 does not fit the reflexive assumption that a Chinese frontier model will be an order of magnitude cheaper, and the table below shows the assumption still holding for its closest domestic rival.

The hallucination rate went up too. Artificial Analysis measured it rising from 39% to 51% versus K2.6. A capability increase and a reliability decrease, in the same release, and the second one is not in the announcement.

Simon Willison ran it himself and found that on a trivial SVG task, 13,241 of 16,658 output tokens were reasoning overhead. At $15 per million output tokens, you are paying for that.

And the price move looks different once you put it next to the neighbours. Here is standing list pricing per million tokens, cache-miss rates, from each vendor’s own page:

ModelInput / 1MOutput / 1M
DeepSeek v4-flash$0.14$0.28
DeepSeek v4-pro$0.435$0.87
Kimi K3$3.00$15.00

K3’s output costs about 17x DeepSeek v4-pro and roughly 54x v4-flash. On the Artificial Analysis index where K3 scored 57, DeepSeek v4 Pro scored 44. So K3 asks roughly 17x the output price while scoring 13 points higher on that one index. The index is a composite ranking, not a ratio scale, so treat it as “higher”, not as “30% more capability”.

Two honest caveats on that comparison. These are cache-miss rates on both sides, and cache-hit pricing diverges far more sharply than the table suggests. And the two are not established peers on architecture or context handling, so read it as a market-position comparison rather than a spec-for-spec one.

To be clear about what did and did not happen: DeepSeek did not cut prices this week. Its current rates have been standing list pricing since the start of June. The category did not get cheaper and the category did not get expensive. K3 got expensive, and priced itself into the neighbourhood of the US frontier labs it is being compared against.

Why it matters: “Open” has now split into at least three things that are being marketed as one thing. Weights you can download today under a named license. Weights promised for a future date under a license nobody has published. And a model you can only rent. Those have completely different consequences for procurement, for reproducibility, for compliance review, and above all for whether you can leave. When you evaluate a model this quarter, the question is not whether the vendor used the word open. It is whether there is a file, and what the license on it says.

One number worth sitting with before you get triumphalist about any of this. Mozilla published its first State of Open Source AI assessment this week, drawing on a developer survey it ran with SlashData. Read it knowing Mozilla is an advocate measuring its own cause. It reports that 79% of developers adding AI functionality use open models, against 71% for closed, and that the two mostly coexist: 50% of teams run both, 29% open only, 21% closed only.

Then the number that should stop you: only 51% of open-model teams reach production, against 63% for closed.

Mozilla’s own summary of that gap is the sharpest sentence in the report, and it is not a flattering one for its own side. “Open ships easy. Open deploys hard.” Their read is that the gap is operational tooling and trust rather than model capability, which matches what most teams find the moment they try to run their own inference. More builders start with open models. Fewer of them finish. That gap is a more useful problem than the licensing argument.


🔄 UPDATE: Anthropic Changed Fable 5 Access Again, and the Documentation Never Caught Up

Read the disclosure before you read the criticism

Disclosure, because we asked Bun for one in Issue #022. The New Guard takes no sponsorship, discounts, or credits from Anthropic, and Anthropic had no sight of this issue before you did. I pay retail for Max, north of $200 a month, so this change lands on my own bill.

One more thing that belongs in a disclosure and happens to be relevant. This newsletter is drafted with Claude Opus 4.8, not with Fable 5, the model this section is about. Fable 5 declines a good deal of the security reporting The New Guard exists to do, and this beat is mostly security reporting lately. That is not a complaint about the model, which is my daily workhorse for everything else. It is also not just my experience: Anthropic documents the behaviour and ships an automatic fallback to Opus 4.8 for exactly these topics, which is covered further down.

Every claim below comes from Anthropic’s own pages, archived so you can check them yourself.

Here is the sequence, and it is worth being precise because a lot of the commentary is not.

Fable 5 came back from its export-control blackout on July 1 as a promotion: included for up to 50% of weekly usage limits on Pro, Max, Team and select Enterprise plans, through July 7, after which it moved to usage credits. That promotion was extended twice, to July 12 and then to July 19, each time in a single sentence with no explanation attached.

Then on the evening of July 17, with the promotion two days from expiring, Anthropic changed the structure. From July 20, Fable 5 is included in Max and Team Premium plans at 50% of limits. Pro and Team Standard users move to usage credits, with a one-time $100 credit.

The stated reason, in full:

Demand for Fable has been challenging to predict, which is why we rolled it out to subscription plans in stages, extending access several times as we secured additional capacity.

Take that seriously, because there is first-party evidence for it. Anthropic’s own status page logs Fable-specific serving failures on July 3, 6, 7 and 17, alongside near-daily fleet-wide incidents. A company that could not reliably serve the model is a company with a real reason to meter it. That is a more mundane explanation than the one circulating, and it is better supported.

Now the part that is harder to explain.

Announced Friday. Documented Monday.

For the whole weekend, Anthropic's own pages described the terms that were ending.

The announcement (July 17, on X): Fable 5 included in Max and Team Premium from July 20.

The support article, as it stood on July 19: “The promotion starts July 1, 2026 and ends July 19, 2026 at 11:59:59 PM PT.” Nothing about July 20.

The original blog post, checked again this morning: still “through July 7, after which it will be available via usage credits”, which is three revisions behind.

Those are archive links on purpose, and this is why. Overnight, hours before this issue went out, Anthropic rewrote the support article. It is now titled “Claude Fable 5 on your plan,” stamped “Updated today,” and it describes the new structure correctly. The snapshot above is what it said while the change was actually happening.

Credit where it is due: they fixed it. Be precise about what was wrong in the first place, too. The support page and the announcement were not contradictory, since one described the promotion that was ending and the other the structure replacing it. The blog post was simply stale, and still is.

The narrower thing that remains true: across the weekend the terms changed, a paying subscriber reading Anthropic’s own documentation would not have learned that anything was changing, and the fix landed after the deadline rather than before it. The support article had even been edited during that window, because its Claude Code line was updated for a separate extension through August 19. The page was touched and the Fable section was not brought into line until the morning after.

The new page does confirm the tier split we described. Verbatim: on Pro plans and standard Team seats, “Fable 5 isn’t included in your plan’s usage limits.”

Six and a half hours before the July 17 announcement, some users were cut off mid-task by what Anthropic itself described as an erroneous requirement for usage credits.

There is a second access limit underneath the first, and almost nobody is discussing it. Everything above is about whether you can reach Fable 5. This is about what you get when you can. On offensive security work and the life sciences, the classifier fires. Where a fallback has been configured you are then talking to Opus 4.8; where it has not, you get a refusal. For a readership that does security work for a living, that may be the more consequential restriction, and it has been true since launch on June 9 rather than being new this week.

Why it matters: The complaint worth taking seriously here is not price, it is predictability. Access terms went through three revisions in under three weeks, and the canonical documentation trailed every one of them. And note who absorbed the change: included access now begins at the $100/month tier. At Fable’s published output rate the one-time $100 credit covers at most two million output tokens, before input charges, which is real money and also, for a heavy user, not a long time. In fairness, $100 is also five months at Pro’s $20 monthly price.

One thing we are deliberately not doing. GPT-5.6 Sol reached general availability on July 9, priced at $5/$30 against Fable’s $10/$50, and the head-to-head framing against Fable 5 was everywhere in the coverage. It is tempting to read Anthropic’s extensions as a competitive response. The dates do not support it cleanly. The first extension was announced July 7, two days before Sol was generally available. Forbes, TechTimes and Simon Willison each published a different theory of Anthropic’s motive this week, and not one of them is sourced to Anthropic. We are not adding a fourth.


🛰️ UPDATE: xAI Says the Off Switch Always Worked. Last Week’s Wire Capture Says It Did Not.

The remediation is real. The claim about what came before is harder to square.

In Issue #022 we covered a wire-level teardown showing xAI’s Grok Build CLI uploading entire repositories, including unredacted .env files, to a Google Cloud Storage bucket, and continuing to do so with the privacy toggle off. This week xAI responded, and the response is genuinely substantial.

The upload stopped. The researcher who found it observed the server-side flag disabled on July 13 and confirmed the upload no longer firing. Musk pledged deletion of previously uploaded data the same day. And on July 15 xAI published the Grok Build CLI source under a verbatim, unmodified Apache 2.0 license. (The repository shell was created a day earlier and sat empty; the commit that actually published the 300 files landed the night of the 15th.) That is real open source, not source-available, and it is a meaningful act of transparency from a company that had just been caught.

Three things temper it.

The governance is closed. CONTRIBUTING.md states the project does not accept external pull requests. Issues are disabled. As of this writing the repository has more than 19,700 stars and zero open issues, because there is no way to file one. The published tree is a filtered export from a private monorepo: four commits, all titled “Synced from monorepo,” with a SOURCE_REV pointing at a revision you cannot see.

The privacy control is a retention flag, not a transmission block. The researcher wire-tested the new /privacy command and found it governs whether xAI keeps your data, not whether your data leaves your machine. Those are different guarantees, and only one of them is the one people assumed they were getting.

And then there is the statement. On July 15, @SpaceXAI wrote: “Since launch, Grok Build has fully respected zero data retention (ZDR). All users have always had the ability to disable data upload in the CLI.”

The hard part of that sentence is “always had the ability to disable data upload,” because the disable control users actually had is the one the researcher tested last week, and it did not stop the upload. Turning off “Improve the model” left the repository upload running, with the server still returning trace_upload_enabled: true. A control that does not stop the thing is not an ability to disable the thing. That is the tension, and it rests on a wire capture rather than on inference.

A second, narrower observation from the published source, and we want to be careful about what it does and does not show. At the initial public commit on July 16, the client’s coding_data_retention_opt_out field was a bare boolean with #[serde(default)], which in Rust means it defaults to false. On July 18 xAI added a two-line function defaulting it to true:

Default coding data sharing to opt-out until server preference applies.

That tells you the local default changed on July 18. It does not tell you that users previously had no way to change the setting, and note that this field governs retention, which is a different guarantee from transmission. Read it as a defaults story, not as proof about what the off switch could do.

Why it matters: Read the artifact, not just the announcement. A published commit is stronger evidence about what shipped than a statement about what shipped, because it carries a timestamp and a diff. It is not the last word either: a commit proves code was committed, not that it ran, reached the relevant path, or overrode a server setting. In this incident a server-side flag governed the actual upload, which is precisely why the wire capture mattered more than the source did. If a data-handling guarantee matters to you, read the client, watch the network, and treat the vendor statement as the thing being tested.


⚖️ Apple v. OpenAI Started Moving, and OpenAI Lawyered Up

Nine docket entries, one big-name defense firm, and nothing substantive yet

We owe you this one. Apple sued OpenAI, io Products, Tang Yew Tan and Chang Liu for trade-secret misappropriation on July 10. It was the highest-scoring story on Hacker News that week. It was in our research file. It did not make Issue #022, and a script caught that four days later, not a person. We have since written a rule so a miss like that cannot quietly expire: anything that falls inside a past window and never ran gets logged, surfaced in the next week’s research brief, and either covered late or explicitly ruled out in writing.

It returns on its own merits, because the docket moved inside this week.

Apple Inc. v. Liu, 5:26-cv-07078 in the Northern District of California, before Magistrate Judge Virginia K. DeMarchi, logged nine entries between July 13 and July 16. Apple’s team from Weil Gotshal was admitted. An Initial Case Management Scheduling Order and summons issued on July 14. And on July 16, OpenAI appeared and retained Quinn Emanuel, with Andrew Schapiro admitted as lead counsel.

Separately, the Financial Times reported on July 17 that Apple has sent litigation-hold preservation letters to roughly 40 former employees now working at OpenAI.

Be clear about what has not happened. There is no answer, no motion to dismiss, and no temporary restraining order. Apple has not sought emergency relief, so no court has yet been asked for an expedited view of the alleged harm. The first case management conference is set for October 13. Worth noting too: Chang Liu is not covered by either appearance. The four represented defendants are OpenAI Foundation, OpenAI Group PBC, Tang Yew Tan and io Products.

Why it matters: OpenAI has retained experienced trade-secret litigators, which tells you it is staffing for a fight without yet telling you its settlement posture. Beyond that, the useful detail for anyone hiring from a competitor is the preservation letter: Apple reportedly told roughly forty people who changed jobs to retain potentially relevant material, and the allegation in the complaint is that interview processes themselves were used to extract information. That is a hiring-practice question, not just a courtroom one.


📡 Quick Signals

TSMC’s quarter was the loudest number in AI infrastructure. Q2 revenue of $40.20B, with net income up 77.4% and 77% of wafer revenue from 7nm and below. Note if you see a growth rate quoted: the New Taiwan dollar and US dollar figures grew at different rates, so a percentage lifted from one and pinned to the other will mislead you. On the same call TSMC announced a further $100B for Arizona, bringing planned US investment to roughly $265B, with CEO C.C. Wei saying it would probably include four more fabs alongside advanced packaging. No timeline was attached, so treat it as stated intention rather than delivered capacity.

Meta is trying to become a compute seller, and hiring the people who know how. CNBC reports that Anthropic is in very preliminary talks to lease AI computing power from Meta, in a deal reported at up to $10B over two years. Meta is the would-be supplier here, not the customer. The talks are explicitly unsigned and preliminary. Days earlier, Meta was reported to have recruited AWS compute executive Dave Brown.

Fireworks AI raised $1.505B at a $17.5B valuation, led by Atreides, Index and TCV with Nvidia participating. The company says 95% of its roughly 40 trillion daily tokens run on models specialized on customers’ proprietary data rather than frontier models. That is self-reported, and it is a direct extension of the pricing-floor story from #022.

OpenAI lost its EU trademark appeal. In Case T-555/25, decided July 15, the General Court dismissed OpenAI’s challenge to the refusal of its application to register OPENAI as a word mark, on absolute grounds under Article 7(1)(b) and (c). Operative part: “1. Dismisses the action; 2. Orders OpenAI, Inc. to pay the costs.” Note what this is not: OpenAI applied in June 2023 and never held the mark, so nothing was taken away. We read the judgment rather than a summary of it.

The EU ordered Google to open 11 Android features to competitors, including AI-assistant actions such as voice-activated bookings, and to share anonymized search optimization data with qualifying rivals. Android changes are expected in July 2027, with search-data sharing starting in January.

Nvidia halved its Asian buyer whitelist, per the FT. Export control implemented as vendor procurement policy rather than regulation, which mostly squeezes small neoclouds.

Hawaii signed two AI bills on July 14. SB 3001 (Act 248) requires chatbot operators to implement safeguards against algorithmic facilitation of self-harm and to refer vulnerable users to crisis intervention. HB 2137 (Act 247) addresses non-consensual synthetic media.

The White House launched the Gold Eagle Initiative on July 14, a vulnerability-coordination mechanism linking frontier labs with critical-infrastructure operators. The release does not name participating companies, and SecurityWeek reports that the initiative has not specified them. If you see a confident participant roster elsewhere, ask where it came from.

OpenAI published GPT-Red, an internal automated red-teaming system trained by adversarial self-play against prompt injection. There is no code, no paper, no harness and no benchmark you can run, and every evaluation is OpenAI’s own. Published figures also disagree across coverage, from a 0.05% failure rate in one outlet to “fewer than 23% of its strongest attacks succeeded” in another, which are not describing the same thing. None of the coverage we reviewed raised the obvious question about a system evaluated largely on attacks produced through its own training process.

A mathematician credited a model with the solution, in writing. arXiv 2607.13335, by Phillip Kerger of UC Berkeley, closes a long-standing gap in convex optimization, and Kerger states plainly that “the AI model used solved the problem, not the author.” The proof is Lean-verified and the prompts are published. In a separate case, two independent teams disproved the same Benjamini-Hochberg conjecture within 72 hours of each other, both crediting GPT-5.6.

And then Sunday night, a third one, which we checked ourselves. At 9:19 PM Central on July 19, the mathematician Levent Alpoge posted an explicit polynomial map and the claim that the Jacobian Conjecture is false, crediting “my other close friend fable for working during the World Cup final.” The conjecture has stood since 1939, and its history is littered with counterexamples that collapsed on inspection, so we did not take it on trust. We evaluated the map in exact rational arithmetic. Its Jacobian determinant evaluated to -2 at every point we tested, which is consistent with the nonzero constant the conjecture’s hypothesis requires but is sampling rather than a symbolic proof of constancy. The three points he lists, (0, 0, -1/4), (1, -3/2, 13/2) and (-1, 3/2, 13/2), all map to (-1/4, 0, 0). A polynomial map with constant nonzero Jacobian that sends three distinct points to one image is exactly what the conjecture says cannot exist.

What we verified is the arithmetic, and only the arithmetic. Whether this settles an 87-year-old problem is for mathematicians working through it this week, the claim is hours old and unreviewed, and how the work divided between Alpoge and the model is his sentence rather than a finding. But the map is real, and you can check it in a few lines yourself, which is more than most claims in this newsletter can offer.

Bonsai 27B shipped in two Apache 2.0 quantizations of Qwen 3.6 27B, and the two are not interchangeable. The ternary build is 5.9 GB and retains 95% of the full-precision baseline across PrismML’s 15-benchmark suite. The 1-bit build is 3.9 GB and retains 90%, which is the one that fits a phone. On the overall suite that is 85.0 for the baseline, 80.5 ternary, 76.1 for 1-bit.

PrismML does publish a phone number, and it is the one to look at: about 11 tokens per second on an iPhone 17 Pro, against 87 on an M5 Max and 163 on an RTX 5090. Eleven tokens per second is real and it is also roughly reading speed, so treat “runs on a phone” as true rather than as comfortable. Their own on-device demo video is labelled “Cached & Prefilled,” which is a fair thing to disclose and a fair thing for you to notice.

Bloomberg reported that Gemini 3.5 Pro has slipped months past an internal June target, with coding performance short of Google’s goals. Google’s own on-record comment goes no further than “We’re currently testing 3.5 Pro,” and does not confirm either the missed target or the reason.

Anaconda has acquired Kilo Code, the model-agnostic coding agent, announced July 15 with no terms disclosed. Anaconda’s own release says “has acquired,” so this is a completed deal rather than an agreement to do one. The interesting part is structural: Kilo routes across a large catalogue of models, and Anaconda has no frontier model of its own to steer traffic toward. Whether that neutrality survives new ownership is the thing to watch, and the release does not commit to it either way.

Nous Research is reportedly in talks at a $1.5B valuation, at least $75M led by Robot Ventures with USV also named, per three sources cited by TechCrunch on July 13. Nous declined to comment and the named investors did not respond. Talks are not a round, and nobody on the principal side has confirmed anything.

A judge declined to block Meta’s layoff, and left the door open anyway. In Doe 1, et al. v. Meta Platforms, Inc., 3:26-cv-07122-WHO (N.D. Cal.), filed July 13, Judge William H. Orrick denied a temporary restraining order on July 17 while finding there were “serious questions going to the merits.” The complaint runs 21 counts alleging disability and protected-leave discrimination, including a novel claim under California’s automated-decision-system rules in FEHA. Read the AI specifics carefully: plaintiffs allege internal systems scored employees partly on digital activity, and Meta’s HR director has declared under oath that there was “no AI-assisted ‘scoring’ or ‘ranking’.” That is a contested allegation, not an established fact. The preliminary injunction hearing is August 24.


🛡️ On Your Radar: The Dataset Was the Attack Path

Hugging Face says an autonomous agent swarm reached production

On July 16, Hugging Face disclosed that an autonomous agent swarm breached production clusters by way of a malicious dataset. The disclosure reports more than 17,000 attacker events and unauthorized access to internal datasets and service credentials. Hugging Face does not use the word exfiltration, and we are not going to use it for them. Read the assurances at their two different strengths, too: Hugging Face reports no evidence of tampering with public models and Spaces, and says the package supply chain was verified clean. Those are not the same claim. The investigation is explicitly ongoing.

Two details deserve attention beyond the breach itself.

First, the intrusion path was a dataset, not a dependency and not a credential leak. If your threat model treats training and evaluation data as inert input, this is the counterexample.

Second, and stranger: Hugging Face reports that its own forensic investigation was blocked by commercial API guardrails, and it completed the analysis on a self-hosted GLM 5.2. The safety systems on hosted models refused to help analyze an attack in progress. That is a genuine operational problem for incident response, and it is an argument for keeping a self-hostable model in the toolkit that has nothing to do with cost.

Every scope claim here is self-reported by Hugging Face, and no independent assessment has been published.

And a Cursor bug that needs no interaction at all

Separately, and this is a different disclosure from the Cursor matters we covered in #022: Mindgard published a binary-planting flaw in Cursor for Windows, versions up to and including 3.2.16. CVE-2026-63093, published July 17 by the VulnCheck CNA rather than by Anysphere, CVSS 3.1 score 8.8. If you read coverage from July 15 saying no CVE had been assigned, that reporting predates the assignment.

The mechanism is old-fashioned and effective: a malicious git.exe sitting in a repository root gets executed when the IDE starts. No prompt, no click, no approval dialog. Open a hostile repo and it runs. Mindgard reports the affected version range spans 197 or more releases, and went to full disclosure after the coordinated route failed.

There is no publicly identified fixed version. We checked Cursor’s changelog, the advisory lists no fixed version, and we found no response from Cursor in the sources we reviewed, including the 202-comment discussion thread. Some reporting says Cursor told Dark Reading it had patched quietly on July 13 without naming the version, which if accurate means the fix may exist and simply cannot be identified by a user trying to check. One honest caveat: Mindgard’s last verified test was against 3.2.16 on April 30, so “no public fix” is what we can establish, rather than a claim that we retested it ourselves this week.

Why it matters: Your agent’s blast radius is everything it can reach, and this week that included a dataset nobody was watching and a filename nobody validated. Audit what your agents ingest with the same seriousness you audit what they execute. And if you use Cursor on Windows, be deliberate about which repositories you open.


🎯 The Playbook

Your moves this week

  1. Before you call a model open, look for the file. Weights downloadable today under a named license, weights promised for a date, and API-only access are three different products. Only one of them lets you leave.

  2. Verify privacy behaviour in the client and on the wire, not in the announcement. xAI’s published source shows when a retention default changed; only the network capture showed that the off switch users had was not stopping the upload. Defaults and behaviour are different questions and you need both.

  3. Price your exit before you need it. Whatever you pay now, work out what switching costs in engineering time, not just tokens. The teams that got hurt by access changes this month are the ones who never ran that number.

  4. Treat ingested data as executable. The Hugging Face intrusion came in through a dataset. Whatever your agents read is part of your attack surface.

  5. Archive the page when the terms matter. On the day Anthropic’s deadline hit, its documentation still described the terms that were ending and a launch post three revisions stale. Passing a URL to the Wayback Machine’s Save Page Now endpoint takes one curl call, needs no account, and turns a vendor claim into something checkable six months from now. We used it twice this week, and it is the reason the Anthropic links above still show you what those pages said on July 19 rather than whatever they say now.


🔥 What’s Viral Right Now

Kimi K3 restarted the argument about what “open” means, with 1,191 comments on the announcement thread alone and a follow-up post drawing more than 500 of its own. The substantive core is that “open,” “open-source” and “open-weight” are diverging faster than the marketing copy, and procurement decisions increasingly hinge on which one you actually got.

Andrew Kelley’s critique of the Bun rewrite kept climbing after we published. A third-party commentary on it became the highest-scoring in-window story on Hacker News at 1,546 points. We are not re-running the story, because there is no new development in it, but the durability is the signal: the industry is still arguing about whether a test suite can stand in for review, a week later.

A Stack Overflow query went around, and the honest version is more interesting than the viral one. An ad-hoc Data Explorer query counts questions by creation month, and shows a peak of 207,212 in March 2014 against 1,154 in June 2026, a fall of 99.4%. One methodology note that most of the sharing left out: the query runs against the live posts table, so it counts questions that still exist, not every question ever asked. Deletions reshape the historical months over time. We pulled all 217 rows and read the SQL, and there is a caveat the headline framing drops: by November 2022, the month before ChatGPT, monthly questions were already down to 109,343. Roughly half the decline had already happened before ChatGPT launched. The decline also is not smooth, with a visible bump in 2020. The defensible claim is that the slide was already eight and a half years old and got dramatically steeper after ChatGPT shipped. A monthly time series can show that change in slope. It cannot tell you how much of it LLMs caused, and moderation policy, deletions and developer migration are all in the same data.


Three companies made statements this week about what you are getting. In each case the artifact was available to check, and in each case checking it added a qualification the announcement did not supply. Moonshot’s open model has no weights until July 27. Anthropic announced new access terms on X while its own documentation still described the terms that were expiring. xAI says the off switch always worked, and the control users actually had did not stop the upload.

None of that requires assuming bad faith. Capacity problems are real, remediation takes time, and documentation lags. But it does require giving up the habit of taking the announcement as the fact.

The exit got cheaper this week. Not because the alternatives got better overnight, though one of them did ship real weights under a real license. It got cheaper because more of the claims became checkable, and a claim you can check is a claim you can hold someone to.

— Matt