Regardless of the true cost, it seems that professional mathematicians now need to wary about what they put into a LLM and think hard about how to disclose and publish a result.
This is all but guaranteed now.
Mathematicians/Scientists/Researchers need to stop sharing freely with "AI Companies" and have explicit clauses in place in their publications about not using their research without their explicit consent.
There should be a clear legal distinction between using research data for AI model-training vs. another researcher using it.
Come up with a legal framework, establish procedures for sharing and using others work and have a single scientific body in charge of enforcing it.
Just putting a clause in a publication won't prevent it from being used as training data. Information wants to be free.
The frontier LLM vendors do sell enterprise licenses which contractually guarantee that your prompts won't be used for training. (Maybe they'll secretly violate the agreement but in principle it's legally enforceable.) Scholars and universities who care about credit and attribution will either have to purchase those licenses or run their own private open-weight LLM instances.
I don't like this and I wish it weren't true, but I think the period of "information wants to be free" is coming to an end, it was a relic of a bygone era. Increasingly, making your information free means you're the sucker who is doing free labor for AI companies, or worse, you're helping your competitors. Paywalls, login walls, and rate-limits are going up everywhere: there's the GitLab news on the home page right now, and sites like Twitter, Reddit etc. which used to be publicly-readable are now gated (and Xitter is using the legal system to shut down any bypasses).
I hate this but I don't think there's any going back now that LLMs exist.
"Information wants to be free" never meant that people want to release their information; it meant that information is very hard to keep secret, and that everything leaks like a sieve, and especailly that once it's out, it's out forever.
Exactly. While there are a few academics who work in private for years and then surprise the world with an amazing breakthrough, most of modern science and mathematics is a collaborate process. Researchers make gradual progress on hard problems, and discuss issues with colleagues and students along the way. Some of those collaborators will then pass on the information to social media or public discussion forums or free-tier LLM prompts or whatever and it gets incorporated into the next round of training runs.
Two people can keep a secret if one of them is dead.
Even the $20 tier of ChatGPT has privacy settings that forbid using the user's data to be used for training. The question is, whether this setting is respected.
Would this legal framework cut both ways? When AI companies use AI to make and publish mathematical discoveries, would they be able to legally prevent professional mathematicians from using them?
This is about cooperation before publishing results. And they will keep everything medieval secret, else some big company steals it and claims it their own.
So then the solution is to publish more often (e.g. on a public blog) even if your ideas are not fully developed in order to establish priority and show you are doing something.
Analogously with software development, it's always been good practice to write things down, but since the start of this year it's become dramatically more important for everyday work.
I misinterpreted the article. Its more about sharing ideas within the company, because teams and team members are apt to steal ideas and implement them with AI faster than the originator. I guess this was always theoretically a problem but its especially pronounced now because of the commonality of layoffs, and exacerbated at my company due to the failing stock price.
Stop treating AI companies and their software as somehow unconstrained, above-the-law actors. It's delusional that anyone buys that. Regulate them appropriately.
At the same time, mathematicians should be using sophisticated, specialized LLM tools in much more sophisticated ways than lay people. There should be no way lay people can compete. There are new tools to master and if you use your slide rule, you won't keep up. It's a chance for mathematics productivity to boom.
With apologies to Baudelaire: The greatest trick exploitative powers ever played was convincing the people that it was impossible to imagine anything else.
Sometimes blinkering people so that they never ask the question "why is this being done to me" or "why is justice not possible" is much easier than finding an answer that will get them to go away.
Like we can prevent the rest of our society from devolving into the Medieval Era of secrecy: by treating individuals with respect and dignity and not as the ore from which resources can be profitably extracted.
Those were fought for, with large strikes and sadly bloodshed and conflict.
The pressing issue is that in the industrial revolution, capital needed labour so strikes are ineffective. How do people fight for their rights in a system that sees no need for them? Especially in one where power is concentrating and politics is frequently on sale to the highest bidder.
Why is socialism a pipe dream? Unions and co-ops aren't pipe dreams, and if we had nation wide unions or converted most businesses into co-ops, we would arguably be living under a socialist economy. If people had direct voting power over economic and business issues that would be socialism.
Some methods are more realistic than others, but I don't see a requirement for any outlandish ideas.
Prior to LLMs there was some minimal effort required to snipe someone and possibly your reputation was attached otherwise there would have been no point in publishing to begin with.
Now anyone with a few dollars can do it, many who don't have a reputation to worry about.
It's the same problem as YouTube AI slop, AI-generated music, and everything else. Don't you dare tweet or blog about a video idea or hum a few bars from a song - within an hour 27 people will have posted AI slop rip-offs. There's something uniquely depressing about being beaten to the punch by a thief that doesn't deliver the same soul-crushing impact as having someone copy you after the fact.
I welcome the era of secrecy. After living so much in this era of open information where everyone seems to know everything, secrets may be a way to make things more interesting again.
There is a financial incentive to selling access to the leading LLMs needed to find whatever secret result there is out there. See the play station hypervisor 0day from the other day. If a LLM can find someone's secret 0day they are flaunting around they can find a math proof someone else says they have.
wouldn't it be dystopian secrecy cause a flock camera is equally capable of watching the mathematicians as it is the public citizens.
Maybe mathematicians are smart enough to never ever buy such piece of shit on higher principle, regardless of their actual fiasco?
I am a (former) mathematician and know many more.
They are not. Mathematicians are as human as most of the rest of us here.
This is all but guaranteed now.
Mathematicians/Scientists/Researchers need to stop sharing freely with "AI Companies" and have explicit clauses in place in their publications about not using their research without their explicit consent.
There should be a clear legal distinction between using research data for AI model-training vs. another researcher using it.
Come up with a legal framework, establish procedures for sharing and using others work and have a single scientific body in charge of enforcing it.
The USA doesn't have legal frameworks any more, you just buy and sell the right to do what you want. Even our supreme court is disingenuous now.
Just putting a clause in a publication won't prevent it from being used as training data. Information wants to be free.
The frontier LLM vendors do sell enterprise licenses which contractually guarantee that your prompts won't be used for training. (Maybe they'll secretly violate the agreement but in principle it's legally enforceable.) Scholars and universities who care about credit and attribution will either have to purchase those licenses or run their own private open-weight LLM instances.
I don't like this and I wish it weren't true, but I think the period of "information wants to be free" is coming to an end, it was a relic of a bygone era. Increasingly, making your information free means you're the sucker who is doing free labor for AI companies, or worse, you're helping your competitors. Paywalls, login walls, and rate-limits are going up everywhere: there's the GitLab news on the home page right now, and sites like Twitter, Reddit etc. which used to be publicly-readable are now gated (and Xitter is using the legal system to shut down any bypasses).
I hate this but I don't think there's any going back now that LLMs exist.
"Information wants to be free" never meant that people want to release their information; it meant that information is very hard to keep secret, and that everything leaks like a sieve, and especailly that once it's out, it's out forever.
Exactly. While there are a few academics who work in private for years and then surprise the world with an amazing breakthrough, most of modern science and mathematics is a collaborate process. Researchers make gradual progress on hard problems, and discuss issues with colleagues and students along the way. Some of those collaborators will then pass on the information to social media or public discussion forums or free-tier LLM prompts or whatever and it gets incorporated into the next round of training runs.
Two people can keep a secret if one of them is dead.
Even the $20 tier of ChatGPT has privacy settings that forbid using the user's data to be used for training. The question is, whether this setting is respected.
the existence of the triplets NSA/CIA/GRU implies an imperitive no.
Would this legal framework cut both ways? When AI companies use AI to make and publish mathematical discoveries, would they be able to legally prevent professional mathematicians from using them?
I'd guess universities might starting hosting open source models. They can probably actually afford to, unlike individual mathematicians.
Though maybe if there's a flurry of math-optimized agents coming up, like there are small coding agents, those might be feasible to host personally.
Simple: if you don't publish your work, we don't fund you. Why is this even a question?
This is about cooperation before publishing results. And they will keep everything medieval secret, else some big company steals it and claims it their own.
So then the solution is to publish more often (e.g. on a public blog) even if your ideas are not fully developed in order to establish priority and show you are doing something.
Analogously with software development, it's always been good practice to write things down, but since the start of this year it's become dramatically more important for everyday work.
I don’t follow… if you publish your underdeveloped ideas, these companies will develop them for you, which is the problem
Similar principle applies to open source or any other creative endeavor put in the public domain.
anyone else already experiencing this in a corporate environment? I know I am. Its not just between teams either, its within them.
The age of a Warhammer 40k Tech Priest is upon us! Praise to the Machine God.
Just, without the aliens, space travel, or mech-suits.
Got half the other downsides though!
More on the upcoming priesthood and monasteries: https://news.ycombinator.com/item?id=49743400
Would you expand on your experiences? Can you not share ideas now? What if you enhance them with an LLM before sharing them?
I misinterpreted the article. Its more about sharing ideas within the company, because teams and team members are apt to steal ideas and implement them with AI faster than the originator. I guess this was always theoretically a problem but its especially pronounced now because of the commonality of layoffs, and exacerbated at my company due to the failing stock price.
I would see stuff like this even before AI.
Stop treating AI companies and their software as somehow unconstrained, above-the-law actors. It's delusional that anyone buys that. Regulate them appropriately.
At the same time, mathematicians should be using sophisticated, specialized LLM tools in much more sophisticated ways than lay people. There should be no way lay people can compete. There are new tools to master and if you use your slide rule, you won't keep up. It's a chance for mathematics productivity to boom.
With apologies to Baudelaire: The greatest trick exploitative powers ever played was convincing the people that it was impossible to imagine anything else.
Sometimes blinkering people so that they never ask the question "why is this being done to me" or "why is justice not possible" is much easier than finding an answer that will get them to go away.
Like we can prevent the rest of our society from devolving into the Medieval Era of secrecy: by treating individuals with respect and dignity and not as the ore from which resources can be profitably extracted.
Any solutions to this issue that aren't unrealistic socialist pipe dreams?
Only realistic capitalist nightmares.
You could've asked that about people saying we want 8 hour work days, sick leave, and safer work 100 years ago, and yet we got them.
Those were fought for, with large strikes and sadly bloodshed and conflict. The pressing issue is that in the industrial revolution, capital needed labour so strikes are ineffective. How do people fight for their rights in a system that sees no need for them? Especially in one where power is concentrating and politics is frequently on sale to the highest bidder.
Why is socialism a pipe dream? Unions and co-ops aren't pipe dreams, and if we had nation wide unions or converted most businesses into co-ops, we would arguably be living under a socialist economy. If people had direct voting power over economic and business issues that would be socialism.
Some methods are more realistic than others, but I don't see a requirement for any outlandish ideas.
Might as well go tell depressed people to just stop being depressed
Or homeless people to just buy a house.
you don't worry about it for a couple years and check back to see if its a real problem or a panic induced engagement generator
We don't.
Prior to LLMs there was some minimal effort required to snipe someone and possibly your reputation was attached otherwise there would have been no point in publishing to begin with.
Now anyone with a few dollars can do it, many who don't have a reputation to worry about.
It's the same problem as YouTube AI slop, AI-generated music, and everything else. Don't you dare tweet or blog about a video idea or hum a few bars from a song - within an hour 27 people will have posted AI slop rip-offs. There's something uniquely depressing about being beaten to the punch by a thief that doesn't deliver the same soul-crushing impact as having someone copy you after the fact.
I welcome the era of secrecy. After living so much in this era of open information where everyone seems to know everything, secrets may be a way to make things more interesting again.
Maybe wise ones will return to paper.
There is a financial incentive to selling access to the leading LLMs needed to find whatever secret result there is out there. See the play station hypervisor 0day from the other day. If a LLM can find someone's secret 0day they are flaunting around they can find a math proof someone else says they have.