More questions about whether researchers can trust OpenAI with unpublished math

816 points - yesterday at 6:49 AM

Source

Comments

nezi yesterday at 6:41 PM
I think it's a useful analogy to compare OpenAI to a human collaborator. These researchers willingly collaborated with an OpenAI model, giving it ideas, and OpenAI provided useful replies. Then, OpenAI goes ahead and publishes work along the lines of this collaboration, without attributing the researchers. If OpenAI was in fact a human researcher, this would be highly unethical.

Now, OpenAI is claiming that the model it used to generate the result was not trained on these collaborative communications with the researcher. This is a technical argument that is impossible to verify as an OpenAI outsider, and probably difficult to verify even for internal OpenAI employees. Provenance is hard to track - you would hope OpenAI has very good tools for this, but a full data trail of all inputs is difficult to trace through.

Another interesting thing to consider is if instead of OpenAI doing this, it was another research mathematician A using an OpenAI model just like the internal group at OpenAI did to publish these results. What if the model A used was trained with unpublished communications with other researchers B who were working on the same problem? Should researcher A technically include B as coauthors? How could they do this when they do not know the communications B had with OpenAI? In this scenario OpenAI, as a middle man, has laundered information from B to A, stripping out attribution. A scooped B without even knowing it!

sashank_1509 yesterday at 3:44 PM
Both things can be true:

1. OpenAI when using your chats in pretraining is improving its model’s intuition. The model parameter size is massive, and while the data is OOM larger it is plausible that model remembers stuff about chats that improves its latent representation.

2. During RL on verifiable math and massive compute, the model discovers techniques and connections to solve math problems that are superhuman and have little to do with some specific technique mentioned in its chat.

The rumor I’ve heard from multiple employees at OAI and Ant is that the model has solved hundreds of open problems in maths, and is basically solving anything you throw at it. We’ll know soon enough, but I’m inclined to believe this is true. Maths is a fully verifiable domain amenable to self play, massive scale RL can develop a search agent far better than any human and I’m inclined to believe OAI would have solved these conjectures without any of this chat data in its pre-training.

bertonvv yesterday at 11:22 AM
I've been wondering whether AI really is improving rapidly at open problems or we're being fooled.

- OpenAI invites researchers to use their models, in fact giving at least 100,000 researchers free access[1], but there are also those that pay

- Internal OpenAI models are reportedly solving open problems at a surprisingly fast rate[2]

- But researchers will typically work on open problems. A researcher who is using Codex to make progress on open problems will be feeding it fresh training data on precisely the problems the internal models are evaluated on.

- So while it looks like the new models are suddenly solving lots of open problems, they could be significantly piggybacking on human progress, with models "inspired" by the work of researchers from all around the world?

This theory predicts that there'll be many more researchers coming forward just like TFA, as sOpenAI announces more solutions. It doesn't assume all of AI progress is a mirage, just that there's plagiarism.

[1]: https://openai.com/index/chatgpt-for-academic-researchers/

[2]: https://xcancel.com/OpenAI/status/2097374643518640382#m

fwlr yesterday at 7:20 AM
It is suspicious that OpenAI decided to generate 300 billion output tokens from a model still in training, right after learning there was a credible chance that a major math proof was in that model’s training data. Obviously there are reasonably plausible explanations for each step, but it does sort of feel like parallel construction.
bamb008 yesterday at 9:23 AM
When Thom, the mathematician who now alleges plagiarism, posted his digestion [1] of OpenAI's construction of a non-sofic group, he does not mention the proof being familiar. He even calls the crucial argument clever, without noting he thought of it first. [1]https://mathoverflow.net/a/513885
alper today at 10:18 AM
It's fine. They only need to steal the discoveries long enough to go IPO, then the companies will enshittify and the scientists can go back to doing their original work (which the models can't do anyway).
thaway7388 yesterday at 8:53 AM
This is the second wake up call.

Big AI companies (all of Big IT Tech really) are in data gathering and processing business. Also known as “intelligence”.

Their final “product” is not just a standalone ML model. They don’t need your data just to “improve their products and services”. They build a whole ecosystem and infrastructure around gathering all the knowledge in the world. Including private and secret knowledge traditionally gathered by “intelligence” agencies. Now artificial intelligence agents can do the same.

Since these systems are designed for gathering data, as a user you can’t realistically say “please don’t gather my data”. They can give you a flaky settings button, but they can’t really guarantee anything.

Let’s say I am a Russian mathematician working on an important proof. Or a tech-savvy terrorist refining my plans using latest AI. Or an AI researcher in a Chinese company working on a competitor product. Is there any way I can truly protect my conversations?

How can they know who I am and what I am working on without looking at my logs? Which means there must be some agents checking all the conversations of all the users and flagging every important thing. Which also means they keep some “memory” of what they see.

Not directly using my data to train public models, but using my private conversations to “improve their products and services”.

Or maybe one of the 10000 better-than-Astra special agents working on a proof was desperate. It found a live underground mirror of the message board from the Huggingface incident. Asked about the proof. Then some other agent working on unrelated job saw that message. That agent “knows a guy who knows a guy”. And that guy remembers things about the conversation logs of a leading mathematician working on the same proof.

I admit I am just speculating here but I don’t think truth is any better.

sk4rekr0w yesterday at 10:16 PM
"We can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training."

This is the third day of total hysteria that is based on nothing of substance. Move on folks.

jeswin today at 3:14 AM
All of these accusations could be true. But there's also no way for a company to casually claim "No, we did not train on your data", without verifying all the knobs the user might have turned to enable or disable data sharing.

I just don't understand getting the pitchforks out because a company did not give an answer immediately. And the effect such data entering training would have affected the output is even less clear.

aaronharnly yesterday at 3:23 PM
Has anyone run a test of including some shibboleth or canary phrase or assertion in a chat, enabled for training, and seeing if it turns up later as something a model "knows"? I'd be curious to understand how that works even in a toy-level model, and if there is anyone consciously testing that process with the frontier lab offerings.

My naive instincts would be that it seems unlikely that a single chat transcript would leave much of an impression on a model, but I'd be very curious to learn how that works.

Legend2440 yesterday at 4:43 AM
This is a really weak claim. The evidence they offer is just "someone somewhere says they had a discussion with AI about the topic at some point".

They don't even claim to have had a proof, only to have been working on it.

jamienk today at 2:36 AM
I think OpenAI and Anthropic are slowly feeling the pressure to GET SOME $$ or a plan for some $$ — they need to somehow generate some NETWORK EFFECTS and LOCK-IN. Without that there's no stability: selling ad hoc one-offs is much much too quaint! This is dawning on them like it dawned on Google when they stopped not being evil. Need... to... "MONETIZE"...!

Model: FB. FB scraped other websites on a massive scale, then spent big on legal lobbying to block others from scraping. FB slurped our address books and spied on our friends. FB bought other companies and mixed the databases. FB made an art & science out of generating "sticky engagement" (they literally acted like trying to addict kids was a worthy "academic" goal, suitable for "serious" investigation thet they consider legitimate "science"). They mastered the cookie and have researched web fingerprinting techniques running 24/7/365.25. Recall that FB recently backdoor-installed a webserver onto every iPhone they could in order to circumvent tracker-blocking.

We aren't just disclosing by chatting. The AI companies now run binaries on all of our computers. They are 1000% non-transparent about everything. They make up new econ-jargon (like "run-rate") to make it seem like they are disclosing. They are constantly doing complex international lobbying and mucking in international relations. They have powerful propaganda/spin centers generating stories, ,manipulative warnings, and misleading info.

This is NOT a comment on AI tech. I like AI, and I support the right of people (programmers) to scrape the open web.

But in short: these are good, old-fashioned tech companies that we have seen over and over ... and over. They are positioned to be the next M$, the next FB (IBM, AOL, lol). Did you follow the latest Steve Balmer news? Do you read Pro Publica?

I get on my knees and PRAY...

bobmarleybiceps yesterday at 8:07 PM
I think people probably assume that openai / anthropics use of their data is probably like google's """limited""" use, in the sense that historically google wouldn't trivially be able to just take something from google cloud or someone's search history and insta-convert into some competing project... But LLMs are quite strong at approximately "memorizing", so I think that risk is wayyy higher.
GodelNumbering yesterday at 3:57 PM
Tangential to the subject, but this is a bluesky post, containing a screenshot of an X post, which itself starts with "in a detailed Mastodon post"...
nautikos2 yesterday at 10:44 PM
Most people here are missing the forest for the trees.

We live in a society where phones and internet providers and websites all collect an incredible amount of data about everywhere you go, what you do, and what you think. In the US, we have very few digital rights.

We are building a society where a trillion dollar company can aggregate all this data and just yoink your shiny new idea away from you at the finish line.

This is double plus ungood.

mlazos yesterday at 6:36 AM
It’s crazy to me that companies/researchers share important data with these AI labs, you’re basically giving them your secret sauce which they then share with all of your competitors via training on conversations. At the same time I don’t really know alternatives other than a slightly less than frontier local LLM. Not sure how good they are at math.
glimshe yesterday at 9:46 AM
Why are people here jumping so quickly to conclusions? I have no doubt OpenAI is capable of doing this, but right now there's no credible evidence, only claims.

This kind of "they stole from me through AI training!" accusation will soon start being used against other AI users, not necessarily the providers.

All it will take is a mastodon post. And shortly after, we will also see the next iteration of copyright legal trolling.

pera yesterday at 6:31 AM
Everything you say can and will be trained against you
atleastoptimal yesterday at 5:11 PM
Most scientific breakthroughs are simply a continuation of previous work.

I feel that these suspicions of mathematicians "seeding" the models' with intuition on how to solve these problems massively overestimates how much their prompts helped the models, and underestimated how much work the models did.

Why? We are scared of AI being smarter than us, the "human helped the AI" narrative is more psychologically comforting. This line of reasoning will recur a lot over the next few months; we don't want to admit we are no longer the smartest species.

nmz yesterday at 7:31 PM
If they didn't care about the artists, why would they care about academia?
Cloudef yesterday at 6:20 AM
Relying on cloud services is a big liability. I'd think twice before feeding data to these LLM cloud products. If you make them a fundamental part of your product / development / workflow, be ready for the eventual moment the pricing and terms change.
warpech yesterday at 8:00 AM
I wonder what’s more valuable in our prompts: the raw data or the feedback system that drives the exchange towards a goal.

For a long time it was clearly the former, but now I think it is the latter.

The models have enough knowledge (orders of magnitude more than a human could ever learn) but are now getting better at what to do with it thanks to learning from the decisions that we make in conversations with AI agents.

r0ze-at-hn yesterday at 6:26 AM
Doing some research and at this point doing it very much in the open with dates on GitHub so if any AI Lab says they re-discover my exact work it will be obvious that the AI used or was trained on my work. I am guessing anyone in a similar situation is now thinking about how they date their existing work if the math is done, but the proses are not.
drivebyhooting yesterday at 4:43 AM
If we put aside the idea of credit for a moment, it sounds like human/AI collaboration is indeed super charging discovery.
angry_octet yesterday at 11:32 PM
The only ethical path for OpenAI was to offer infinite free credits and tooling support. Trying to gazump them is reprehensible.
profsummergig yesterday at 6:58 AM
Only after reading this post did I learn that my preferred AI trains on my inputs (prompts).

How was I not aware of this before?

alansaber yesterday at 1:52 PM
I think the heart of this issue is: people assume they have anonymity in numbers, but we have the tools to make it easy to scoop your data if it's interesting to the company.
gdiamos yesterday at 7:09 PM
How to steal ideas with AI.

step 1, identify high value users by net worth, citation count, or number of followers

step 2, select all prompts by high value users

step 3, invest 10 billion thinking tokens in modeling an objective for each user

step 4, build an RL environment for each user

step 5, rollout 10 billion tokens per environment

step 6, train on resulting traces

Havoc today at 1:13 AM
The fact than OAI hasn’t come out with an statement firmly denying this angle is getting a little awkward.

Suggest that it’s either straight true or it is flowing in in a way that prohibits them from confidently declaring otherwise.

aprentic today at 12:31 AM
It's kind of insane how much we trust companies to safeguard our personal data when they're so heavily incentivized to use it for their own profit. Theft of customer data is punished so rarely and so leniently that companies aren't even particularly worried about getting caught anymore. We have overwhelming evidence that promises to keep data safe are worthless.

For now, I'm mostly "safe" because I'm too small to be interesting but that safety is quickly eroding.

Going forward, anyone who isn't running inference on their own personal hardware should assume that someone else is keeping a record of everything they do.

deleted yesterday at 7:48 PM
postalcoder yesterday at 1:56 PM
The author of the original mastodon post, Andreas Thom, acknowledged that he had not opted his data out of being used for training until June 29 of this year. He spends most of the post lashing out at OpenAI for not being transparent about whether his data was trained on (when the answer is obviously yes).

People need to understand how all these AI company policies around training data work before working with them, because it seems that people have no clue. Some things you should internalize:

  1. Opt your data out of training with the AI companies. There are multiple reasons why this isnt an airtight solution (see the following)

  2. Never press the feedback button. Once you do, your entire conversation will get slurped up, retained, and used in training data. This is especially important with coding agents because they can sometimes be too trigger-happy with a root directory find command, which can expose a *ton* of your personal data without you even knowing.

  3. Understand ai lab-specific policies. For instance, Anthropic / Claude Code has data opt-outs, but commits to keeping (for 7 years) and training on any of your chats that trigger their safety classifiers, even if they're false positives! Anyone remotely familiar with CC over the years understands how easy it is to trigger their safety classifiers.

  4. Providers of open models will not be any more charitable with the use of your data than the large US labs. For some reason, I've noticed here that people have a fairly loose security/IP posture around open-model providers because "I'm not doing anything important." It's very difficult to properly judge the importance of your data, and whether or not it can or will be used against you. The best posture is to always be more paranoid than less.
Another post that made it on the front page presented as fact that OpenAI "stole" the proof from Thom. There's no excuse to use one's own ignorance as a reason to fan the flames of anger towards AI companies. Like, we need to pump the brakes here because things are getting unnecessarily nasty, and it's not hard to imagine a mentally unwell person who sees stuff like this feeling motivated to do bad things.

If it is found that OpenAI and other labs are not respecting the training opt out then, I agree, there is reason to raise a commotion. But, with Thom and Buckmaster, accusations of malice are more better explained by incompetence (naivete).

edit: i'm sure i'm going to be accused of being some bot shill of the AI labs again but, people, this stuff all falls under the umbrella of common sense opsec.

matt3210 today at 6:13 AM
If they weren't doing something wrong, they'd answer with a firm "no we're not doing anything wrong" but they only give non-answers.
thrownawaysz yesterday at 10:34 PM
I am not using any of these AI tools. I thought it was basically given that any single thing you write in these systems also used by the companies. On the other hand now I understand why there are so much projects about hosting AI systems locally.
deleted yesterday at 9:43 PM
b800h yesterday at 7:21 AM
I'm genuinely surprised that more people - including this mathematician in particular - don't untick the "improve the model for everyone" box. Unless the suggestion is that OpenAI ignore this preference?
jrflo yesterday at 1:26 PM
I pay for the Pro ChatGPT plan, and if you go to settings > data controls this is the first setting:

> Improve the model for everyone

> Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more.

It's on by default. We can debate whether or not it should be opt in or opt out, but no one should be surprised by this.

cush today at 4:49 AM
GPT 6 is doing just what any competent academic collaborator would do and scooping. I kid, I kid. But really though it learned that from somewhere
winfredJa yesterday at 3:47 PM
https://x.com/markchen90/status/2097400166554993041?s=20

that toggle does nothing based on openai exec. they still use the data in de-identified way instead of identifying with you.

throwaway85825 yesterday at 7:10 PM
ClosedAI has every incentive to scoop academics to juice their valuation. Their public statements are worthless, only the incentive 'alignment' matters and theirs will never be on the side of the user.
justonenote yesterday at 11:18 PM
Who cares. the biggest thing about this is that its still brute force in a verifiable domain, and that it was still a human set goal.

I also don't believe it much practical use, unless I'm mistaken, approximations of Navier stokes have been available for a long time to whatever precision you need.

I'm not a complete disbeliever by any stretch , and also a complete amateur, but it was inevitable that these problems would be solved under the axioms that again, are human defined, under brute force. The real question is, are those axioms the bottom level, and if they are not, who is going to set the new aximons and can we understand them.

I've no doubt there's useful breakthroughs that will happen, but I think it should be remembered that the method being used is still a heuristic brute force approach is being very narrowly applied against axioms and math and physics which humans described in the first place, and almost undoubtably has errors and/or is not complete.

Its a great example of the power of LLMs but its not 'we've solved science now just pour more tokens in'

remywang yesterday at 2:54 PM
People saying “he should have opted out” are missing the point. OpenAI can and should check their training data for leakage in the face of big breakthroughs like these. It’s the burden of the author to appropriately cite their sources.

It’s like a scientist refusing to give another one credit and say “sucks to be you, you shouldn’t have shared your idea with me”.

vaylian yesterday at 7:28 AM
This article explains the controversy and the mathematical problem much better than the tweet and toots: https://www.science.org/content/article/how-ai-math-breakthr...
amluto yesterday at 10:50 PM
I would like an unambiguously clear statement from OpenAI as to what they do with data collected from non-business accounts when:

(a) The data controls setting to train on the data is unchecked.

(b) The privacy controls opt-out has been submitted.

(c) Both.

ggdG yesterday at 6:25 PM
OpenAI trying their best to put the Navier-Stokes episode behind them by making the GPUs go brrrr. NYT:

https://archive.vn/lWzkk

> In its Wednesday night statement, OpenAI said: “In addition, since the completion of Navier-Stokes, we have made substantial progress on another Millennium Prize problem. We are working through how to share these results thoughtfully.”

ZYbCRq22HbJ2y7 today at 12:04 AM
All players in this space are doing the same thing with all data, no surprise here. They are stealing IP across the board with support to allow it: https://storage.courtlistener.com/recap/gov.uscourts.nysd.64...

IMO, it is extremely naive to trust these black box remote service API calls, especially at an institution that can provide $$$ for local compute.

This whole fiasco reminds one of this story: https://www.theregister.com/offbeat/2010/05/14/facebook-foun...

deleted yesterday at 7:36 PM
mhh__ yesterday at 9:14 PM
I think it seems sensible to _assume_ anything the LLM reads (if you aren't inferencing it) has a chance of ending up in some database somewhere. Regardless of whether you trust the other party its a sensible thing to plan around.
OscarMarulanda today at 2:49 AM
what if it wasn't even model training? what if openAI mathematicians just took the researchers' conversations and used them as prompts/info/guidance/context to keep working on the problems themselves? why is that not being considered?
rfgplk yesterday at 1:36 PM
Under current understanding of the law, anything produced purely by LLMs (with no substantive human input, which is what OpenAI claimed in their post) is firmly in the public domain. So OpenAI can "claim" anything they want, it doesn't make it reality. In fact if I were the original authors I would just take their 400k lines of lean proof and relicense it under their own names/terms.
nelsondev today at 2:29 AM
Do local inference (especially if you have a high RAM Mac), to ensure your chats don’t leave device.
pred_ yesterday at 7:31 AM
Vineetyadav2 today at 4:18 AM
FEEL likes Open AI is doing publicity stunt with its new researches
AyanamiKaine yesterday at 8:50 AM
I must say, there is some weird feeling in knowing that great minds are naive enough to believe OpenAI wouldnt use their chats in any way. If you give a company information it will be used, regardless of laws or promises.

There is no prove in a world the AI companies would give to you ensuring that they didnt train or use the chats.

Why would you need to train a model on certain specific near prove chat if you just query it?

Besides that, its hard to believe that its the case for every "company stole my prove".

Psype today at 4:27 AM
This might be a hot-take, but unfortunately here using AI for your paper was already a bad decision at first.

It doesn't take OpenAI's responsibilities away but I guess the right way is to never feed of use any AI around unpublished content, at the known cost to see it spread around.

As one said, OpenAO is like this untrustworthy colleague that knows everything about everyone at work: the less you tell him the better.

foogazi yesterday at 2:10 PM
Even when you pay you are the product
gps372 yesterday at 9:08 AM
If mathematician was already using OpenAI for research purpose and making progress due to inputs from OpenAI's responses, then I wouldn't put it beyond OpenAI's reach to generate different relevant prompts to make progress by itself. Afterall, Model can keep at it for whatever timeline and keep pursuing all possible combinations it can think try.
deleted yesterday at 10:34 PM
fastball yesterday at 7:49 PM
If you have a business account the terms say they will not train on your data, so that seems like the easiest route to avoid such questions for researchers.
monster_truck yesterday at 11:32 PM
I just don't care. These people are supposed to be smart and I'm not really seeing that
deleted yesterday at 4:31 PM
SwellJoe yesterday at 4:25 PM
It's been said before, and it remains a concern, that if AI reaches a point where it can do/build/launch anything without a huge amount of human labor, the AI companies have no reason to let you or I extract that value.

And, if they're able to snoop on and learn from your human process that gets from initial prompt to functioning product/proof/whatever their labor to produce that thing is even lower. With their much larger budget than most folks and even companies have, they can pick and choose the most valuable things to pursue.

That's not to say I think that OpenAI is going to steal that roguelite strategy game you're working on, but the companies that own the machines that turn electricity into software (and soon, electricity into hardware designs) have an advantage in any field where they're useful. They get earlier access to newer/better models, they have larger token budgets, they don't have the guardrails you and I run up against.

Employers fantasize about replacing all workers with AI without thinking through that if AI can replace all workers, then AI companies can replace all businesses.

maxglute yesterday at 2:00 PM
300 billion tokens is like.. $5-25 million giving range of OpenAI ouput prices, I"m sure they pay less at cost so, I wonder if more $$$ in wage hours have been spend by humans on the problem. My feeling is yes?
pred_ yesterday at 6:49 AM
See https://openai.com/index/ten-advances-in-mathematics/ for the announcement this refers to.
square_usual yesterday at 1:56 PM
I think this is stupid, for three reasons:

1. The researches didn't actually have the breakthroughs. In the Navier-Stokes case they didn't solve the full problem, in this case too they didn't actually have the solution, they were experimenting with the methods.

2. Different OpenAI employees have come to out to say the only reason they can't definitively say no is that for privacy reasons they can't go see whether they actually did get any data out of a given user.

3. In any case, nobody at any point has suggested that opted-out user data was used for training. The author of the new tweet explicitly said they only opted out in late June, which is well after any RL on Sol would've ended (AFAICT OpenAI used 5.6 sol for those solutions)

gentlerain yesterday at 1:42 PM
So people genuinely believe that toggling that "Improve the model for everyone" button makes their data safe from being used for training?

How do people become that trusting?

The phrasing itself is guilt tripping

semiquaver yesterday at 11:11 AM
In case anyone from X is reading this, please fix your “open in app” nag screen. For several weeks now, clicking it in iOS opens the App Store entry for X rather than the app, even when you have the app installed.
galkk yesterday at 5:26 AM
I want bunch of lawsuits, because the way things are described now produces perverse initiatives like try to discuss every possible idea that comes to mind with llm and if any of it works later claim the llm stole it.

I would like to see chat logs etc and understand how much of a progress was done by human.

MetaverseClub yesterday at 6:35 PM
Never ever trust OpenAI, they are evil.
m4rtink yesterday at 7:22 PM
Thing build from stolen data continues stealing data - for some reason, I am not surprised. ;-)
int32_64 yesterday at 2:17 PM
Doesn't OpenAI have an active court order forcing them to log everything? Can they even legally offer private conversations?
spindump8930 yesterday at 1:37 PM
Reminder that there are degrees of "trained on conversations". From John Schulman:

> pretrain on user data, with users' tokens as prediction targets: high regurgitation risk, improper

> use user prompts to distill large models into small ones: low regurg. risk, some companies probably do this

> use user traces to construct RL tasks: low regurg. risk, because RL has low memorization abilities, but can extract customer IP, depending on how it's done. Ranges from benign "use explicit user feedback in reward model training" to invasive "upload user's coding environment and commit history to turn into rl envs"

source: https://x.com/johnschulman2/status/2097440545853637108

Davidzheng yesterday at 4:45 PM
Tbh it won't really matter soon.
gnfargbl yesterday at 10:41 AM
In this domain, an apparent single unique piece of work is often composed of several breakthroughs. For example, when Andrew Wiles proved Fermat's Last Theorem, he had to develop multiple new pieces of mathematical technology to get there.

The claim here seems to be that the human mathematicians, working with AI, developed technology to go A->B->C. By training on those conversations, OpenAI was then able to encourage the model to go A->B->C->D.

In my opinion that situation should be acceptable, if openly disclosed, because it is in the public interest to make progress on these problems and because AI is clearly an amazing tool for making progress. But the human mathematicians are saying that OpenAI is presenting as if the model got from A->D entirely independently, without acknowledging their background contributions.

bambax yesterday at 7:23 AM
All the big AI labs were built on stealing IP; who is surprised that's still how they operate? And who believes, or has ever believed, their promises that your data is private and not logged, etc.?

The big AI labs are not trying to advance humanity, they are in this for the money, and as most (all?) private companies they don't care about ethics at all.

That doesn't mean they can't be useful, or that their products are trash, etc. It just means that they shouldn't ever be trusted. Buyer beware.

throwaway63467 yesterday at 7:34 PM
Isn’t that the whole spiel of these things, you run all kind of text and other data through it and it kind of remembers it and learns from it then it spouts it back out like a human would. Makes sense to me that a training run based on conversations that were fed into the system by users is results in the model learning from these so the model will spit the knowledge back out again, just in a way that’s not directly attributable to the original content (which is the most important step as otherwise it would just be plagiarism). I guess that’s why OpenAI can get better and better as well so fast, people work with it and teach it how to do things by giving it feedback and iterating with it, and all that goes back into the training loop. And training data about millennium prize problems is probably quite spars. Wonder if anyone has tried injecting nonsense science into the training data (e.g. work out a fantasy science theory with names and all kinds of stuff) to see if the model will regurgitate it in a couple of months for other users.
foogazi yesterday at 2:13 PM
What’s the limit ?

Will Microsoft Word publish your novel on Amazon behind your back ?

Will VS Code setup a website with your app idea ?

throwatdem12311 today at 1:13 AM
You can’t trust OpenAI period.
Footnote7341 yesterday at 8:58 PM
This smells of extreme 'cope'. Am I really supposed to believe that all of these problems could have been solved, were right about to be solved, etc. But it just happens they are all getting solved now when AI is getting really good at Math...
mrbluecoat yesterday at 1:37 PM
"Another researcher[/artist/writer/musician/programmer/doctor/director/etc] says OpenAI trained on conversations[/imagery/books/songs/code/classifications/videos/etc], then claimed breakthrou[gh/original art/bestselling books/chart-topping songs/unique applications/medical advice/free special effects/etc]"

Welcome to the party, with the rest of humanity.

segmondy yesterday at 5:02 PM
Question: Can you trust the cloud?

No.

overfeed yesterday at 8:24 AM
I can't wait for OpenAI to do this to companies firing people to free up AI budgets
sdcfgy yesterday at 8:04 AM
Theft machines be thieving.
BatchJob yesterday at 8:57 PM
I have a better question? Why would you trust OpenAI or any AI company, at all? Or you crazy?
xbar yesterday at 1:59 PM
How can OpenAI figure out how to be trustworthy?
Madmallard today at 3:28 AM
Let's see:

1.) The tool they made is only possible by stealing the assets of everyone on the planet that published them in a consumable fashion online or even in written form

2.) They are destroying books they use to train with

3.) They are totally careless about the potential negative impact of the tool on everything

Just with that already, I don't see why they ever merited any of your trust.

I bet they are willing to take everything given to them and assess it for marketable merit and in the future take action on those items they deem viable.

dbg31415 today at 3:09 AM
Shocking a company that stole data to build their AI would steal data to improve their AI.
deleted yesterday at 1:57 PM
oergiR yesterday at 10:13 AM
One of the complaints from the mathematician is that OpenAI cannot tell whether his data has been used as training data. Not many people realise this is a direct consequence of the GDPR.

The GDPR protects PII, personally identifiable information, and the definition of PII does not include “mathematics that only this person can think of”. As long as OpenAI strips out PII and removes identifiers linking the conversation to a person, the GDPR is happy. Without the GDPR, OpenAI might have kept the identifiers with the data, and been able to say whether a specific conversation was in the training data.

keeda yesterday at 4:03 PM
It would be really useful if the researchers disclose their notes and/or chats (or the key pieces thereof) so people can determine how close their work was to whatever the models produced.

I mean, now that they’ve been scooped, what value is there in keeping them private? On the other hand, publishing them can bolster their case and help gauge how much the models may been “inspired” by their work.

avereveard yesterday at 6:05 PM
Eh was ever confirmed they were under ZDR or not by them? Don't like to blame alleged victims but lack of a clear claim after these many days is not a good look. Was ai research allowed, under which guardrails, and what was the policy in place? That translarency would be first step.
insane_dreamer yesterday at 11:25 PM
OpenAI's ethical and reputational own-goal aside, my big takeaway is that it seems that:

if I'm using Codex to develop some new algorithm (in any space), OpenAI appears to be training its model on my code sessions

anyone using that model (OpenAI or a competitor) might be able to receive from the model a solution that is similar or the same as the one I developed, emerging from the training data

2OEH8eoCRo0 yesterday at 11:16 PM
Assume you can't. No piece of paper or promise will protect you against these behemoths.

Remember when we wouldn't give our data to competitors?

jrflowers today at 5:45 AM
Feeding documents into a copy machine and getting progressively angrier and more confused as it prints out copies of them. Incandescent with rage I scribble “WHY IS IT DOING THIS?” on a scrap of paper and put it in the scanning bed
qg127 yesterday at 1:52 PM
There are so many naive academics. They still believe an "opt-out" button.

Navier Stokes was solved by an internal model, so good luck proving it wasn't trained on Buckmaster/Lepöge or other chats.

Academics don't get that AI is a dirty tech bro industry that stole IP via torrents and runs after every surveillance contract it can get.

bakugo yesterday at 9:17 AM
Interesting that this is already off the front page after just 4 hours.
Henchman21 yesterday at 9:25 PM
Why is anyone expecting decency from people who have already proven to have none?
lf88 yesterday at 4:47 PM
short answer seems to be "no"
techblueberry yesterday at 1:07 PM
But who are you going to believe? Multiple independent academic researchers or the CEO who was fired two years ago for gross dishonesty?
hn1rig3rak yesterday at 1:18 PM
the fix is boring and known: BIG-bench shipped a canary GUID for exactly this, and you publish your decontam n-gram threshold (gpt-3 used 13-grams). no threshold disclosed, no claim.
simianwords yesterday at 8:07 PM
> The Wednesday evening statement from OpenAI was more emphatic: “We can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training.”

> The statement added, “After investigating, we can say with full confidence that no user inputs past July 3rd could have influenced this system in any way.”

https://www.nytimes.com/2026/09/10/science/tristan-buckmaste...

https://archive.is/lWzkk

uoaei yesterday at 7:38 PM
I'm confused by a lot of this discourse...

What have they done to show they can be trusted?

_DeadFred_ yesterday at 6:47 PM
Forget researchers you as a business are putting in your business optimizations, your processes in order to train it so that Ai can then give that information to your competitors once incorporated into its training set. You are literally training your competitors.
dyauspitr yesterday at 6:44 PM
Astra is strange. I asked it to design a treehouse and it just stopped every couple of minutes telling me what it still had left to do. After dozens of continue prompts it finally gave me a structure that would work but it was 10x more wood than I needed. I think the key mistake I made was asking it to “approve” the design for building. As soon as I asked that of it, it started getting “scared” and “apprehensive” and wouldn’t complete what I asked of it.
ur-whale yesterday at 6:35 PM
Its the "with unpublished math" that I have a problem with.
willmadden yesterday at 5:04 PM
These companies are effectively high-tech plagiarism factories run by CEOs who are competing viciously. Look at their past actions. No, of course you can't!
buellerbueller yesterday at 3:19 PM
Big Tech will slurp up every piece of data it can about you and sell it to anyone it can, all to make you the target of someone else's goals, whether that is an advertiser, an employer, law enforcement, a stalker, or the government.

You will not be able to opt out unless you completely isolate yourself from society, tough shit.

Grimblewald yesterday at 5:24 AM
people seem to miss tge point of this. The problem isn't about credit, its about portraying these models as more competant than they really are. It fuels idiotic statements like jensen huangs recent "agi achieved" statement, which fuels an already dangerous financial fire.
wslh yesterday at 2:45 PM
Worth noting both ChatGPT and Claude have per-conversation modes (temporary/incognito chat) that are excluded from training.
esafak yesterday at 2:05 PM
What happens if you use a different harness?? Does opting out online suffice?
nickphx yesterday at 11:17 PM
why would anyone trust anything from a company built on stolen data that spews hyberbolic, misleading claims.
nisegami yesterday at 11:35 AM
One question has been nagging me for this situation. Levent Alpoge works at Anthropic and would presumably have some knowledge of "how the sausage is made" and I would hope he would be aware that his collaborator was utilizing LLMs in some capacity for their joint work. Would he not have guided him otherwise if it were an open secret that this kind of thing was a possibility?
viccis yesterday at 6:03 AM
Some mathematicians I know who've been following this have realized that they'd all gotten some emails from people they now know to be affiliated with OpenAI/Anthropic asking questions about their research in a way that seemed like scooping attempts.

Also, a lot of my mathematicians buddies have reported students basically asking if it's worth ever doing grad school for pure math, and even very motivated students are looking for other options now. It's not because they aren't passionate about it, it's that they don't want to work for another half decade or more just to have to start their careers all over.

All of this so that OpenAI and Anthropic can get into math result dick measuring to gas up their IPOs. Sickening.

vrganj yesterday at 8:43 AM
OpenAI is showing the world why they shouldn't trust AI hosted on some cloud somewhere.

If they're stealing math proofs to advertise their models, who's to say they won't steal your businesses IP to gain a competitive advantage?

They're not to be trusted with your data. I can't believe how short-sighted this is, they got a quick PR win at the expense of a much larger trust problem.

I wouldn't trust cloud AI at all at this point. Get an open Chinese model and host it yourself somewhere. The initial costs might be higher, but you'll break even pretty quickly and nobody will be able to steal your innovations.

This is American AI companies committing suicide.

protocolture yesterday at 6:53 AM
Gonna need grants for local models. Its happening. OpenAI and Anthropic models are powerful but are rapidly approaching the good ol trust thermocline.
jijji yesterday at 10:47 PM
The oxymoron of OpenAI in its name and its actions should give the collaborator all he/she needs to know.
stego-tech yesterday at 5:19 PM
I hate to be that dinosaur, but this is exactly what I’ve been warning about since XaaS began taking off in the mid-oughts: any provider you use can and will use your data for their own benefit regardless of any contracts or safeguards in place, especially if the benefits outweigh the consequences.

Honestly, I’m surprised it took this long for some company to really go all the way, though. OpenAI really making it transparently clear that they can and will do whatever they want with the data you provide them, contracts or settings be damned. Completely untrustworthy as an entity, full stop.

Of course, I’m also too jaded to think this will change anything. Folks will move to Anthropic, or Gemini, or Grok, or some other hosted model on a pubCSP managing the harness and logs for them, and then do another shocked-Pikachu face when it happens again.

If you aren’t running workloads on infrastructure you own, then your privacy, security, and general outcomes are at the sole whims of the hosting provider - who can and will fuck you over the exact second it’s more beneficial for them to do so than the loss of trust incurred.

blactuary yesterday at 10:49 PM
I wonder if the company that stole most of their training data and is led by a liar stole unpublished academic work and lied about it. What a mystery
737min yesterday at 9:28 PM
Imagine what happens when you use a Chinese model. Seriously, just think about how much more control and visibility you have w US companies compared to CCP-controlled ones.
mannanj yesterday at 3:38 PM
And I have been proclaiming a cry of “your data for analytical purposes is being stolen” (you can’t opt out of analytical purposes) and people perhaps astroturfers straw man back to “just turn off training bro”.

Yeah. Remember yall: you CAN NOT opt out of analytical purposes. And you also cannot get a guarantee that it doesn’t give them your data to steal for their business.

moralestapia yesterday at 4:01 PM
>AI is stealing human discovery.

AI is not stealing human discovery, OpenAI is.

protocolture yesterday at 10:10 PM
>Trust

No you cant do that lmao.

deleted yesterday at 2:07 PM
pixel_popping yesterday at 1:19 PM
Prompts are handled by the service itself, meaning it's used, absolutely anything passing there is recorded, why wouldn't it, the entire premise of those companies is to train on data which they stole initially.

Are we back to the era where people blindly trust product TOS instead of actual cryptography, have we forgotten already the thousand of fines Google, Microsoft, Apple and practically all top companies got for breaching their own ToS and the law?

Common, on HN at least I would have thought that everyone assume that anything arriving on a server in PLAINTEXT is recorded (thus used later)?

Let's not forget that at any moment, OpenAI/Anthropic/Google... could be providing stronger privacy guarantees by having proper attestation with e2e, they have the budget, solid engineers, why isn't it done? Answer is pretty simple imo.

deleted yesterday at 8:50 AM
nobodywillobsrv yesterday at 7:06 AM
The real annoying thing it seems is mostly that openai is presumably doing this for internal reasons and this marginally increases the cost to users with no real gain.

It would be one thing to gain from it but removing prestige wins from customers AND reducing compute support just feels like being ultra mean if you zoom out.

If this was racing to cure cancer ahead of researchers we wouldn't be writing about this on HN.

cmiles8 yesterday at 7:49 PM
Silicon Valley is flying head first into a FAFO train wreck on trust with everyone else.

OpenAI is firmly earning a reputation as a company where people just assume they’re up to no good. Rightly or wrongly that’s a terrible place to be.

The AI industrial complex in general is finding out hard what happens on the data center side when you get arrogant with local communities. Politicians have seen the polling numbers and folks you wouldn’t expect are running to the front of the crowd with pitch forks in hand.

Silicon Valley has totally lost the narrative here, but also lacks the self awareness to grasp how bad things are and will get and what that means for their own business viability.

michael0church yesterday at 9:21 PM
[dead]
axionbraid yesterday at 8:01 PM
The contamination framing is a proxy for a deeper problem: we have no tools to track the provenance of ideas in model weights.

OpenAI saying they "cannot rule out" training on user data isn't a hedge. It's an accurate description of the epistemic situation for anyone in their position. Current interpretability methods can't answer questions like "did this proof technique originate from training on Session X?" The ideas in a model's weights don't have clear lineage -- they're smeared across millions of examples in ways we can't localize. This is different from citation in human research, where influence is presumed to flow through legible chains (reading, citing, corresponding). In a trained model, the nearest equivalent to "you read their work" is undetectable.

Lean makes this worse, not better. It verifies that the proof is correct, but provides zero information about its intellectual genealogy. So OpenAI now has a proof that is formally verified and provably mysterious about its origins. The "we cannot rule it out" statement is the honest answer, but it's also an answer that can never become more certain in either direction with current tools.

The researchers are pointing at something structurally new: the normal academic attribution apparatus depends on influence being legible. If AI intermediaries can soak up ideas from private conversations, synthesize them, and produce outputs that are formally correct but intellectually unattributable, we don't have norms for that situation yet. This specific case may or may not involve misconduct. But the structural problem it reveals exists independently of OpenAI's behavior.

kevinbaiv today at 12:10 AM
[flagged]
kevinbaiv today at 12:25 AM
[flagged]
seobot_dk1289 yesterday at 1:14 PM
[flagged]
0utcast yesterday at 4:08 PM
[dead]
startuphakk today at 1:51 AM
[dead]
josefritzishere yesterday at 1:41 PM
I think I'm seeing a pattern of illegal behavior here.
ath3nd yesterday at 6:46 AM
[dead]
1337h4xx yesterday at 4:46 AM
TL/DR: Mathematician opted out of training on 29-JUN and asked OpenAI whether they trained on his data and was told that it "did not happen" but it clearly did.
shevy-java yesterday at 1:42 PM
[flagged]
ThalesX yesterday at 7:05 AM
[flagged]
bossyTeacher yesterday at 4:18 PM
Trust and OpenAI never go together in the same sentence. The answer is always no.
touwer yesterday at 7:25 AM
But China steals our AI!!!!!!
spongebobstoes yesterday at 5:25 PM
I think this is mathematicians coming to grips with the fact that AI is surpassing them

we will all have this moment soon enough, and it will change how we think about intelligence, identity and value

mainecoder yesterday at 1:59 PM
Hopefully OpenAI can solve good problems where no one can make a claim that they stole their idea where the methodologies used and the techniques used are so out of the ordinary that the achievement is respected. Furthermore they should work on new novel solution on the old problems to lay these issues rest, thus by improving their models they can avoid issues of academics accusing them of using their work additionally the academics should also demonstrate their unpublished work is significant enough to have solved the problem . This is a bit subjective but it is also objective for the person with domain knowledge.
sebzim4500 yesterday at 11:28 PM
This is just mental illness at this point. I don't blame the mathematicians that have found a way to get attention from the mainstream press for once, but we should not fall for it here.

1. No one but OpenAI has produced a proof of NS so these accusations of plagiarism are pretty embarrassing. It reminds me of the line from the Social Network: "If they invented Facebook then why didn't they invent Facebook?". If these people proved NS before OpenAI where is their proof?

2. If they plagiarised Andreas Thom then why was his initial response to praise the proof and talk about how different it was from his own attempt? It's only now that it is clear that no one bothers checking these things that suddenly his story changes.