I agree. The frontier models are based on training data from tons of copyrighted work. Some of that work was obtained illegally, even. They could not exist without strip-mining the commons. The labs have no moral or ethical ownership to the end result, and others should feel free to treat any company-imposed restrictions on their use as invalid.
I don't expect Tan's position to be based on any kind of real moral high ground, but his conclusion is correct.
I love the "illicit distillation attacks" framing from the incumbents. There's nothing illicit. There's no attack. You just don't like it because it threatens your market position and business model.
If the leading private labs attempt to use the government to pull up the ladder under the pretense of "safety" then the response of the people should be to take such questions out of private hands and nationalize the leading labs.
Or they could abide by the precedents they set and learn to compete. They shouldn't be allowed to have it both ways.
With the recent Navier-Stokes controversy, I think there's a credible suspicion that all your IP you run through these models will end up in these companies' possession. OpenAI themselves has admitted a weak version of this (that prompts might inadvertedly end up improving the model). We don't know the extent of this.
Obviously it's not possible to run a company whose value is predicated on its IP that uploads said IP to a third party which might get access to it.
This could mean every potential serious customer would have no option but to seek alternatives to these online services.
ZDR is based on the exact same pinky-promise as traning opt-outs. There is no technical barrier to OpenAI, or whoever is running your compute, retaining your prompt after they run inference on their servers. If you don't control the hardware the model is being inferenced on, you don't control your data.
Abolish copyright and make it less ridiculous. Sampling music was never a thing that required royalties until the 1990s when I guess someone got angry that rappers were making money off their sampled music. Its insane to me. Make it illegal to transfer ownership of copyrighted work too, only the spouse or one single inheritor who isnt a company can have the rights transferred, after both die, the work enters public domain.
LLMs should just pay a flat fee to use a specific book and thats it. Fees should be reasonable (not a million dollars per book), so long as the model doesnt spit out the entire book.
One of the most infamous legal challenges to sampled music was MARRS "Pump Up the Volume" in the 1980s, and that was preceded by other famous cases. Not sure why you think that started in the 1990s.
All correct, just help me get over the idea of an open-weight Mythos where one or a dozen of us eight billion does something stupid on the bioweapon front. Smart people who’ve exhausted possibilities for what they can do with books and web search and today’s Kimi/GLM.
Figure we’ll have to reckon with this next year in any case, guess we’ll see.
> He also notes that the proprietary AI labs didn’t ask permission when they vacuumed up as much human knowledge as they could to train their models.
I think this should desactivate the moral high ground from which Anthropic is trying to speak. That they would want to make distillation orderly IMHO is fair, but to make it illegal is very rich from any AI frontier lab, really.
Also, as said elsewhere: "Lab" is rich here, for outfits that, facing these giant, energy swallowing black boxes have really no clue what's going on inside.-
The moniker gives them an air of scientific, knowledgeable, tranquil, pro-social, pro bono work.-
Of course they are entitled to kill off a few mice, or pillage the commons to forward their "lab" work.-
“We don’t know what’s going on” is essentially marketing. Sure we don’t _know_ but we have intuitions about why, where, and how to make certain changes…
We also associate laboratories with evil scientists and Frankenstein and the like. I can just hear Boris Karloff (er Bobby Picket) uttering “I was working in the lab late one night. When my eyes beheld an eerie sight… … … …the monster mash”. If anything, I associate _uncertainty_ with labs. The result is never known up front, they’re a place of discovery.
But I get your meaning. What should they be called instead? AI Sausage Factories maybe (cue Upton Sinclair?)?
Both labs even explicitly promise the customer owns the outputs. It feels like they want to have their cake (ensure enterprises don't get spooked away from using as many LLMs as possible) while eating it too (still arguing some level of control over the outputs).
> Ownership of content. As between you and OpenAI, and to the extent permitted by applicable law, you (a) retain your ownership rights in Input and (b) own the Output. We hereby assign to you all our right, title, and interest, if any, in and to Output.
> As between the parties and to the extent permitted by applicable law, Anthropic agrees that Customer (a) retains all rights to its Inputs, and (b) owns its Outputs. Anthropic disclaims any rights it receives to the Customer Content under these Terms. Subject to Customer’s compliance with these Terms, Anthropic hereby assigns to Customer its right, title and interest (if any) in and to Outputs.
> Both labs even explicitly promise the customer owns the outputs.
> to the extent permitted by applicable law, you (a) retain your ownership rights in Input and (b) own the Output
If the argument is that the model itself is under copyright protection then "as permitted by applicable law" would be doing some heavy lifting. Assuming that were true, given that locally-run LLMs exist, what would be illegal: the distillation itself or the provision of service of the distilled model?
YC does better if its startups get open weight frontier benefits. Garry’s just advocating for his book, which is his job. Consider how much capital YC portfolio companies would have to burn until liquidity if they have to pay OpenAI and Anthropic, versus relying on open weight frontier capabilities.
Society as a whole has paid into this technology: through the theft of its intellectual property, by having to deal with the pillaging of so many commons (digital or otherwise) by it, through skyrocketing energy and computing device prices, and even just ordinary investment. Democratize the technology! Or at the very least, don't step in legally to prevent this from happening.
I think OpenAI and Anthropic will go bust, or at least be scrapped for parts in the next 5 years or so. It's clear that the extreme cost used up for training is impossible to recoup, as inference is already being subsidized.
It's also clear that, as Tan indicates, open-weight models will be (and basically already are) just as good as frontier models. It's all about the harness, baby. We will have two main forks in the road, and two new industries created:
- AI hardware (NVidia/Cerebras/etc.), the equivalent of Intel/AMD
- AI software (harnesses, assistants, etc.) the equivalent of Microsoft/Apple
We already saw a glimmer of this with popularity of OpenClaw—the problem is that it's janky, hard to set up, inconsistent, and very hacker-esque. Imo "AI labs" will be a dying breed because there's no real money in the actual models if they get commoditized, which they already kind of are.
Agreed! Allow US companies to innovate by creating an ecosystem of smaller, more efficient open weight models and it will be a net benefit for everyone. Distillation is a good thing.
Preventing token-consumers from developing competing products should be litigated as anti-competitive behavior.
Distilled models are worse than the original, so you can’t fully compete. Also, if all frontier labs did that, there would be nothing left to distill from.
I see it as analogous to companies building fiber in the public ROW during the last big infrastructure bubble. Under the Telecoms Act, these companies had to allow competitors to use their fiber at a fair price.
Similarly, AI companies should be required to allow distillation at a fair price. Fair Use doesn’t make sense as a social contract if it only cuts one way!
But if they tried to set a fair price they would have to report how much money they are losing on each token sold. This might be bad for the real business of ai firms, hoovering up as much capital as they can
Garry Tan and Sam Altman recently did this interview together. They seemed pretty friendly with each other during it. Wonder what Sam Altman would say about Tan advocating for OpenAI’s models to be distilled.
Then again this is the same OpenAI that has gotten into legal trouble recently regarding Apple’s IP so who knows
This is all based on the delusion that Chinese labs are mindlessly distilling the frontier.
I would love for a US lab to be at or near the frontier with an open weight model, but it’s going to take some serious elbow grease, and yes some distillation (which btw OAI, anthropic et al, also use distillation of other’s outputs in their training)
Gates probably honestly believes in UBI; the guy is practical to a fault but evil misleading genius he is not. I actually don’t see any better options than UBI long term.
Ok, but how do the economics of this work? Based on its settlement, Anthropic paid an average of $3000 per work they scanned based on their settlement (https://tech-insider.org/au/anthropic-copyright-settlement-2...). They and OpenAI pay billions per year for a mix of experts and normal people to label or create data. Why would they continue doing this if the value of this is immediately copied by open models? If your goal is to end the economics of generating and buying data for AI (and I recognize for some people this is really the goal) then sure, but if you want AI for various subfields of interest to continue improving then it's not workable.
Back when people made arguments for software privacy, the argument was usually "big business will still pay and consumers wouldn't have paid anyways so it's ok for us to pirate" - I actually think that was fine for business software but terrible for indie games, whose market was 0% businesses.
But in the AI case, it's not like they get to keep some of the value of their investment - it all gets cloned into models that businesses and consumers alike are happy to use. If someone knows how labs could continue to fund data creation and acquisition in this model, please do share!
They can’t they’re literally fucked, and it’s not society’s problem! The whole world doesn't have to bend over to make sure a couple of lunatics who believe they are building a doomsday weapon also have a viable business model
Surely if you hoover up every book in existence to feed into an ai model you must be extracting more than 1.5B in value. If not then it’s not a viable business.
I agree. The frontier models are based on training data from tons of copyrighted work. Some of that work was obtained illegally, even. They could not exist without strip-mining the commons. The labs have no moral or ethical ownership to the end result, and others should feel free to treat any company-imposed restrictions on their use as invalid.
I don't expect Tan's position to be based on any kind of real moral high ground, but his conclusion is correct.
I love the "illicit distillation attacks" framing from the incumbents. There's nothing illicit. There's no attack. You just don't like it because it threatens your market position and business model.
If the leading private labs attempt to use the government to pull up the ladder under the pretense of "safety" then the response of the people should be to take such questions out of private hands and nationalize the leading labs.
Or they could abide by the precedents they set and learn to compete. They shouldn't be allowed to have it both ways.
With the recent Navier-Stokes controversy, I think there's a credible suspicion that all your IP you run through these models will end up in these companies' possession. OpenAI themselves has admitted a weak version of this (that prompts might inadvertedly end up improving the model). We don't know the extent of this.
Obviously it's not possible to run a company whose value is predicated on its IP that uploads said IP to a third party which might get access to it.
This could mean every potential serious customer would have no option but to seek alternatives to these online services.
Almost every serious customer is already using ZDR where nothing is retained at all, instead of "anonymized" data.
They already trained on pirated content, what makes you think they are going to honor ZDR?
Where nothing is retained at all, allegedly.
ZDR is based on the exact same pinky-promise as traning opt-outs. There is no technical barrier to OpenAI, or whoever is running your compute, retaining your prompt after they run inference on their servers. If you don't control the hardware the model is being inferenced on, you don't control your data.
This. Distillation “attacks” are a made up concept. It's as if I claimed that Anthropic made a “training attack” when training on my internet writing.
Abolish copyright and make it less ridiculous. Sampling music was never a thing that required royalties until the 1990s when I guess someone got angry that rappers were making money off their sampled music. Its insane to me. Make it illegal to transfer ownership of copyrighted work too, only the spouse or one single inheritor who isnt a company can have the rights transferred, after both die, the work enters public domain.
LLMs should just pay a flat fee to use a specific book and thats it. Fees should be reasonable (not a million dollars per book), so long as the model doesnt spit out the entire book.
One of the most infamous legal challenges to sampled music was MARRS "Pump Up the Volume" in the 1980s, and that was preceded by other famous cases. Not sure why you think that started in the 1990s.
All correct, just help me get over the idea of an open-weight Mythos where one or a dozen of us eight billion does something stupid on the bioweapon front. Smart people who’ve exhausted possibilities for what they can do with books and web search and today’s Kimi/GLM.
Figure we’ll have to reckon with this next year in any case, guess we’ll see.
"Bioweapon" information is not useful without a lab for synthesis.
Someone with that lab could almost certainly figure out how do something stupid or destructive on their own, or bypass model safeguards somehow.
why can't I use the tokens i paid for anyway?
I'm sure they put some BS in their TOS
I'm also certain that they violated countless ToS when they scrapped the internet for training purpose.
> He also notes that the proprietary AI labs didn’t ask permission when they vacuumed up as much human knowledge as they could to train their models.
I think this should desactivate the moral high ground from which Anthropic is trying to speak. That they would want to make distillation orderly IMHO is fair, but to make it illegal is very rich from any AI frontier lab, really.
Also, as said elsewhere: "Lab" is rich here, for outfits that, facing these giant, energy swallowing black boxes have really no clue what's going on inside.-
The moniker gives them an air of scientific, knowledgeable, tranquil, pro-social, pro bono work.-
Of course they are entitled to kill off a few mice, or pillage the commons to forward their "lab" work.-
> The moniker gives them an air of scientific, knowledgeable, tranquil, pro-social, pro bono work.
The Atlantic argued this (rather well, IMO) a week or so ago - "There’s No Such Thing as an AI ‘Lab’" - https://www.theatlantic.com/technology/2026/09/stop-calling-...
“We don’t know what’s going on” is essentially marketing. Sure we don’t _know_ but we have intuitions about why, where, and how to make certain changes…
We also associate laboratories with evil scientists and Frankenstein and the like. I can just hear Boris Karloff (er Bobby Picket) uttering “I was working in the lab late one night. When my eyes beheld an eerie sight… … … …the monster mash”. If anything, I associate _uncertainty_ with labs. The result is never known up front, they’re a place of discovery.
But I get your meaning. What should they be called instead? AI Sausage Factories maybe (cue Upton Sinclair?)?
> What should they be called instead? AI Sausage Factories maybe (cue Upton Sinclair?)?
That's actually great? Slaughterhouses killing off the collective genius of humanity and grinding it into a bland paste for mass consumption.
I reflected on this myself recently. Model distillation seems to be at least as fair a use as distilling a book.
More than fair if you consider that the tokens are paid for.
Both labs even explicitly promise the customer owns the outputs. It feels like they want to have their cake (ensure enterprises don't get spooked away from using as many LLMs as possible) while eating it too (still arguing some level of control over the outputs).
> Ownership of content. As between you and OpenAI, and to the extent permitted by applicable law, you (a) retain your ownership rights in Input and (b) own the Output. We hereby assign to you all our right, title, and interest, if any, in and to Output.
https://openai.com/policies/terms-of-use/
> As between the parties and to the extent permitted by applicable law, Anthropic agrees that Customer (a) retains all rights to its Inputs, and (b) owns its Outputs. Anthropic disclaims any rights it receives to the Customer Content under these Terms. Subject to Customer’s compliance with these Terms, Anthropic hereby assigns to Customer its right, title and interest (if any) in and to Outputs.
https://www.anthropic.com/legal/commercial-terms
Obviously there is some bad behavior going on in the distillation scene with gray-market token resellers but that is "just" normal fraud.
> Both labs even explicitly promise the customer owns the outputs.
> to the extent permitted by applicable law, you (a) retain your ownership rights in Input and (b) own the Output
If the argument is that the model itself is under copyright protection then "as permitted by applicable law" would be doing some heavy lifting. Assuming that were true, given that locally-run LLMs exist, what would be illegal: the distillation itself or the provision of service of the distilled model?
YC does better if its startups get open weight frontier benefits. Garry’s just advocating for his book, which is his job. Consider how much capital YC portfolio companies would have to burn until liquidity if they have to pay OpenAI and Anthropic, versus relying on open weight frontier capabilities.
If someone proposes the right thing for selfish reasons, do we call that bad? Or do we call it proper incentive alignment?
It's an observation, how you feel about it is a personal opinion. I have no opinion on whether it is "bad vs "good." It just is.
Society as a whole has paid into this technology: through the theft of its intellectual property, by having to deal with the pillaging of so many commons (digital or otherwise) by it, through skyrocketing energy and computing device prices, and even just ordinary investment. Democratize the technology! Or at the very least, don't step in legally to prevent this from happening.
I think OpenAI and Anthropic will go bust, or at least be scrapped for parts in the next 5 years or so. It's clear that the extreme cost used up for training is impossible to recoup, as inference is already being subsidized.
It's also clear that, as Tan indicates, open-weight models will be (and basically already are) just as good as frontier models. It's all about the harness, baby. We will have two main forks in the road, and two new industries created:
We already saw a glimmer of this with popularity of OpenClaw—the problem is that it's janky, hard to set up, inconsistent, and very hacker-esque. Imo "AI labs" will be a dying breed because there's no real money in the actual models if they get commoditized, which they already kind of are.Agreed! Allow US companies to innovate by creating an ecosystem of smaller, more efficient open weight models and it will be a net benefit for everyone. Distillation is a good thing.
Preventing token-consumers from developing competing products should be litigated as anti-competitive behavior.
https://archive.ph/BnceE
If it were so easy why aren't the frontier labs doing it themselves?
Distilled models are worse than the original, so you can’t fully compete. Also, if all frontier labs did that, there would be nothing left to distill from.
I see it as analogous to companies building fiber in the public ROW during the last big infrastructure bubble. Under the Telecoms Act, these companies had to allow competitors to use their fiber at a fair price.
Similarly, AI companies should be required to allow distillation at a fair price. Fair Use doesn’t make sense as a social contract if it only cuts one way!
But if they tried to set a fair price they would have to report how much money they are losing on each token sold. This might be bad for the real business of ai firms, hoovering up as much capital as they can
Freefire
https://youtu.be/ZIaOBAjvc38
Garry Tan and Sam Altman recently did this interview together. They seemed pretty friendly with each other during it. Wonder what Sam Altman would say about Tan advocating for OpenAI’s models to be distilled.
Then again this is the same OpenAI that has gotten into legal trouble recently regarding Apple’s IP so who knows
This is all based on the delusion that Chinese labs are mindlessly distilling the frontier.
I would love for a US lab to be at or near the frontier with an open weight model, but it’s going to take some serious elbow grease, and yes some distillation (which btw OAI, anthropic et al, also use distillation of other’s outputs in their training)
There were comparisons and Muse Spark is so very similar to Fable / Opus... so...
Garry also goes to Thiels silicon valley church.
Like Gates saying there should be UBI, or Musk saying... well, whatever.
They know it won't happen, so arguing for it is 'effectively free' and purely personal marketing.
A bullshit game played by politicians and wannabes.
Gates probably honestly believes in UBI; the guy is practical to a fault but evil misleading genius he is not. I actually don’t see any better options than UBI long term.
Ok, but how do the economics of this work? Based on its settlement, Anthropic paid an average of $3000 per work they scanned based on their settlement (https://tech-insider.org/au/anthropic-copyright-settlement-2...). They and OpenAI pay billions per year for a mix of experts and normal people to label or create data. Why would they continue doing this if the value of this is immediately copied by open models? If your goal is to end the economics of generating and buying data for AI (and I recognize for some people this is really the goal) then sure, but if you want AI for various subfields of interest to continue improving then it's not workable.
Back when people made arguments for software privacy, the argument was usually "big business will still pay and consumers wouldn't have paid anyways so it's ok for us to pirate" - I actually think that was fine for business software but terrible for indie games, whose market was 0% businesses.
But in the AI case, it's not like they get to keep some of the value of their investment - it all gets cloned into models that businesses and consumers alike are happy to use. If someone knows how labs could continue to fund data creation and acquisition in this model, please do share!
They can’t they’re literally fucked, and it’s not society’s problem! The whole world doesn't have to bend over to make sure a couple of lunatics who believe they are building a doomsday weapon also have a viable business model
> Anthropic paid an average of $3000 per work they scanned based on their settlement
Not sure you get to count breaking the law and getting in trouble in your cost-of-doing-business. That's a little too on the nose.
You're basically arguing that a criminal syndicate must be allowed to continue and we're required to make their business model make sense?
%99 of the startups fail, they are venture backed. Nobody or no market forced them to spend like that. It's all their decisions
Surely if you hoover up every book in existence to feed into an ai model you must be extracting more than 1.5B in value. If not then it’s not a viable business.