Rendered at 06:17:35 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
kiwih 8 hours ago [-]
As an academic and regular submitter to arXiv, this is an eminently sensible policy. I like that it's per-submitter and not per-author, that means that the big labs with large collaborations shouldn't be terribly put out (more authors = more submission budget).
I hope something like this could also be adopted for some of our larger conferences - the absolute limits on co-authorship are what seem to cause the most grumbles.
bryan0 9 hours ago [-]
some highlights:
> In September of 2016, arXiv received 9,869 submissions. In September of 2024, arXiv received 20,569 submissions. This September, arXiv received 40,363 submissions, which in turn generated almost 9,000 support tickets for arXiv staff and moderators.
> arXiv now limits submitters to up to two submissions per calendar month, with a limit of three total active submissions at any given time.
totetsu 4 hours ago [-]
It seems there is a lot of this kind of AI risk. I'm trying to find the right label for it. Something that was a common good, that worked at a human scale, is now made unviable because automation has pushed it beyond what it can handle and still be useful.
lhd1 2 hours ago [-]
several authors including the creator of arxiv, Paul Ginsparg, foretell a future in which we "increasingly rely on status markers such as author
pedigree and institutional affiliation as signals of quality, ironically
counteracting the democratizing effects of LLMs on scientific production."
I think this is a very plausible outcome of a flood of preprints. Because there's only so much time in a day and are we going to spend it on papers by less established people who we do not trust? It's a pity. If the science is good what does it matter who produced it.
In my experience, ArXiv has become too slow to publish. My previous paper took weeks to be approved. It was a recommended path for journal submission (posting to a public repository and submitting a link), and after waiting a long time, I just posted it on hal.science instead, and it was up in a day.
Another one of mine has been on hold for 18 days and counting.
I think it is a great service to the community, but clearly they are having a scalability problem.
bonoboTP 7 hours ago [-]
Arxiv papers are not supposed to signify anything. The only reason that people may derive some kind of career benefits is because Google Scholar indexes it and counts it towards the h-index etc. All that metric-based career-optimization going on is quite broken anyway, so this is fixing things from the wrong end as well. (The better end: stop hiring and promoting people based on Arxiv paper counts, so there is no incentive to flood it)
The hosting cost side I can understand, though they received quite some funding recently, but I assume it's not going towards actually running the site, but to who knows what broader impacts and so on.
Moderation for Arxiv is a silly idea. It already shouldn't be taken as a quality signal that something is able to be up on Arxiv. Obviously they should remove illegal stuff, but beyond that, requiring moderation is a misunderstanding of their role and reason for popularity.
If hosting is too expensive, I guess an alternative aggregator could also arise. With just metadata and a hash of the pdf that can be hosted anywhere, and as the sumbitter, if you move the file, you can change the URL.
The main reason for Arxiv's existence is the timestamping and the easy referencing. (Though I admit that the stable hosting is also a pretty important part, but they mention moderation effort as the reason, not the hosting costs.)
Academia is losing sight of the forest for the trees, can't see more than an arm's length ahead of their noses.
kragen 6 hours ago [-]
> Moderation for Arxiv is a silly idea. It already shouldn't be taken as a quality signal that something is able to be up on Arxiv.
It sounds like you're advocating for arXiv to become viXra. On https://vixra.org/all/2610 the second paper is currently "Correct Interpretation of the Great Discoveries in Particle Physics: I. Reconsidering the Higgs Boson via Vedic Vortex Structure", followed by a paper arguing that classical electrodynamics is bogus, "Gauss’s Flux Theorem Does not Hold in Time-Varying Electric Fields", and then "On the Boundary Problem in the Origins of Matter, Life, and Consciousness".
I don't think arXiv will continue to receive "quite some funding" if this is what it contains.
miki123211 5 hours ago [-]
Vixra suffers from the "Witch Hunt problem" because Arxiv exists.
If Arxiv was the only game in town and it allowed everything, the result would be more like GitHub or Substack. That is, you wouldn't be able to trust everything hosted there, and a lot of it would be low-quality promotional garbage, with filtering moved to a different layer.
You could do some kind of karma system for example, where the prestige of your institution, your own karma, number of citations, papers published in high impact-factor journals etc would affect karma weight. You could have journals be like "awesome-x" lists on Github, with authors who have already been published there able to endorse other papers. I believe the mathematics community is trying that.
bonoboTP 5 hours ago [-]
> You could do some kind of karma system for example, where the prestige of your institution, your own karma, number of citations, papers published in high impact-factor journals etc would affect karma weight.
The story that academia likes to tell itself about itself is that it is skeptical and iconclastic and objectively evaluates each piece of science on its own virtues and it doesn't matter who says it. A wrong thing said by someone famous is still wrong, and truth spoken by a nobody is still truth. Of course reality is a bit further from this ideal. But outright admitting it would be hard.
It's a bit like the "in America, anyone can become president!".
kragen 5 hours ago [-]
The funny thing is that you'd think that arXiv would suffer from the same problems because journals and conferences exist.
bonoboTP 5 hours ago [-]
People who submit to and review for conferences have known for quite some time how random that process is. Also they are slow. In a field that advances leaps and bounds on a monthly time scale, holding a conference half a year later or publishing a journal article more than a year later is just equal to irrelevance. The reason conferences still matter is because of legacy systems that are slow to adapt to the new reality (and because of networking opportunities).
But actually, thinking back, 10 years ago there may have been more of a stigma around a paper that was only on arxiv.
jltsiren 5 hours ago [-]
If your result is not worth publishing in a journal a year or two from now, it't not worth submitting to arXiv today. Just write a blog post about it and move on. And maybe write a paper once you have discovered something of more lasting value.
bonoboTP 5 hours ago [-]
It's worth publishing somewhere because PhD students need to graduate. And there are many PhD students because it's a ticket to industry labs. And because many professorships around popular topics were created. Those people need to graduate, so there have to be papers. Blog posts don't let you graduate. A good arxiv paper that just got rejected for unfair reasons, but otherwise gets a good community reception can still count towards your graduation if you have a sensible advisor.
Not sure if you understand how the sausage gets made. People won't spend effort at scale (there are always exceptions) on stuff that's not incentivized, such as blog posts. If you work on something, the goal is to make it count. But if you fail on that, arXiv is a second best option.
Also, papers that get published at conferences also get out of date pretty fast anyway. That's just how fast moving fields work.
But things may end up moving the way you are saying, as long as incentives keep up.
Journals (1-2 years) -> conferences (6 months) -> arxiv (weeks to write once the work is done, then days to publish) -> next thing (immediate AI writeup, code, results, tests, open scripts, weights, full reproducibility, full access allowed for other AIs to scrape and learn from, and reorganize and recombine on a speed of days or hours from result to adoption)
jltsiren 4 hours ago [-]
It's not been my experience that papers published in CS conferences get outdated quickly. Maybe AI/ML is different, but it's a weird edge case that should not dominate the discussion. Even CS as a whole is an outlier, because most fields that use arXiv don't have conference papers. If the expectations of AI/ML or even the wider CS are too different, maybe they could start a preprint repository of their own.
One of weird aspects of CS research is the tendency to publish small papers. In more experimental fields, a paper may represent years of work for the first author. But CS students often finish multiple projects every year. Maybe it's because of the focus on conferences. You want to attend conferences (to build your networks, and because it's interesting and valuable), but you often don't get reimbursed if you don't have anything to present. Students from other fields can just present their work in progress, but CS students have to come up with publishable results.
Maybe CS should also strive towards bigger projects. Instead of publishing individual results (which someone often improves a year or two later), the papers could focus on the outcomes of entire research directions. That way you are more likely to produce something of lasting value. Something that is worth writing up and rewriting and polishing, until you understand the ideas much better than after the initial write-up.
bonoboTP 6 hours ago [-]
This is always the fate of explicitly "alternative" platforms, because only those people go there who are pushed out of the mainstream, who will tend to have some serious deficiencies, otherwise they'd try to be on the mainstream platform. Same with unmoderated social media turning into far-right places. Once a place obtains this type of reputation, it will be even more repellent for anyone not like that who wants to make it clear they are not like that, leading to a cascade, where the thing becomes a hermetically isolated "radioactive" place.
By the same principle, someone in the early 2000s could reasonably say "only weirdos date online", but the mainstream can shift also.
6 hours ago [-]
contubernio 3 hours ago [-]
Arxiv is the principal forum for publishing mathematics research of quality. Journals serve a purely archival and bureaucratic purpose; almost no value is added by the refereeing process (probably more is lost) and journal publishing neither improves accessibility nor diffusion of results.
bee_rider 6 hours ago [-]
Agree that it basically shouldn’t be necessary. But unfortunately Arxiv isn’t able to change the bad incentive structure that the broader hiring ecosystem has set up, right?
bonoboTP 6 hours ago [-]
If academics and their committees actually deeply considered the scientific qualities of applicants, gaming the numbers would not matter. So I just find it hard to blame AI slop instead of the broken process that academia has settled on.
Part of the reason that the value of Arxiv papers has gone up, is that it's pretty well known that the conference review system is already quite broken and random and lazy and superficial. So being rejected from there is not a good reason to ignore a paper, so conference-rejected Arxiv-only but good papers are a regular occurrence. If conference review had better signal, arxiv could have zero signal.
The other issue is the time delay. Conferences have stopped being an actual place to learn about new work, the way it had been 10+ years ago. Now it's all outdated stuff, and as a specialist you already know the important papers months before the actual conference, so the conference is a networking event basically. The real exchange of ideas moved to Arxiv and Github and social media, with its own problems, hype, algorithmic engagement optimization incentives etc.
But the pressure that shifted away from conferences to arxiv is continuing. The old guard was pearl clutching already by the erosion of conference rubber stamps as this big prestige. The newer ones are now trying to hold back the tide at the Arxiv line. But it will keep on moving faster and faster. And what is going to matter in the end is not how things used to be, but what delivers actual value. If the slop is slop, it will dwindle. If it starts to be actually good, all this will just turn into basically a dock-unions-opposing-automation story.
paulpauper 4 hours ago [-]
Moderation for Arxiv is a silly idea. It already shouldn't be taken as a quality signal that something is able to be up on Arxiv. Obviously they should remove illegal stuff, but beyond that, requiring moderation is a misunderstanding of their role and reason for popularity.
Arxiv papers are regularly cited and assumed to be suitable for peer review. So some standards need to be maintained . it's not a free for all.
bonoboTP 4 hours ago [-]
How about reading before citing? The bar was already very low for getting into Arxiv. If you used that as a quick decider of whether the work is good and should be trusted, you were already doing it wrong.
the__alchemist 9 hours ago [-]
Nice writeup and the policy is a step in the right direction. I suspect simple rate-limit measures like this will be a big improvement.
I suspect a logical conclusion Arxiv and elsewhere may be an identity management system with an aggressive filter, and shared blacklists. I suspect that classifying people as spam/slop-submitters, then banning them (or whatever identity they used; name, email, name + organization etc), applying incremental rate limits over the general one, may be required.
iterance 6 hours ago [-]
If the queue size grows beyond volunteer capacity then volunteer time becomes a scarce resource. In my experience (in other organizations), more likely than out and out block listing is aggressive prioritization. It takes sensible design to determine which papers get to consume moderators' time first.
Then, if someone posts rejected papers regularly and is deprioritized, it naturally follows that they may try to submit three more... and get stuck naturally. There is no need to ban them. They will still get a response, and they will get a fair shake - when there is time.
lhk931122 6 hours ago [-]
There may be great researchers who work on several papers in a month, but even with this rule they can just split them and submit to arxiv, so I think it is very reasonable. AI slop is everywhere now (and also increasing), and I think this fast and forceful action by arxiv is effective at stopping this wave of slop.
paulpauper 4 hours ago [-]
publish everything on zenodo for timestamp/priority and then slowly integrate to arxiv
move6729 9 hours ago [-]
This will just force high SNR Posters to other platforms. ArXiv is rate limiting itself out of relevance.
anvuong 8 hours ago [-]
Any example of high SNR authors who consistently putting more than 2 papers months after months after months for this to be a problem?
AlotOfReading 9 hours ago [-]
It doesn't seem humanly possible to produce >2 "high SNR" papers per month. If you have a larger batch of related papers, it's easy to spread submissions over multiple months. You should probably be doing that anyway, because comments on one paper may affect the others.
NotOscarWilde 9 hours ago [-]
> It doesn't seem humanly possible to produce >2 "high SNR" papers per month.
The situation right now is pretty crazy, for example this established TCS professor [1] has at least 4 arXiv submissions co-authored by him in September [2]. This includes the recent breakthrough on the matroid secretary problem, which has a pretty interesting AI story of its own, if you haven't seen it yet [3].
Of course prolific professors can work around this by working with younger researchers who upload the work, but this just makes the rate limit a solo author bottleneck, which feels a bit weird.
[3]: Concurrent Discovery Disclosure: The proof of the main result in this manuscript was obtained in a conversation with ChatGPT-6 Astra on Tuesday, September 15, 2026 at 1:02 AM PDT. We then prepared this manuscript for public release, with the intent of uploading it on the morning of Thursday, September 17, 2026. In the early morning hours of September 17, while finalizing the submission, we discovered the manuscript of Abdi, Banihashem, Hajiaghayi, and Mittal, uploaded on September 16, 2026, which contains the same result via an essentially identical approach. We are sharing our manuscript nonetheless in case our exposition is of independent utility to the community, and we hope this experience stimulates broader discussion about concurrent discovery in the AI era. (From https://arxiv.org/abs/2609.20797 .)
AlotOfReading 8 hours ago [-]
I don't see what we lose if lead authors are forced to submit work themselves instead of going through the PI's account.
But yes, it'd be nice if arxiv also rate-limited co-author submissions with some higher number to discourage the most flagrant PI co-authorship abuses.
jltsiren 7 hours ago [-]
The situation is pretty crazy, but it just means the bar for publishable results will go up. Theoretical computer science is a field built around conferences. The number of papers that can be published in reputable conferences ultimately depends on the size of the community and the willingness and ability of its members to attend those conferences.
Gregaros 8 hours ago [-]
Insane. No. Mathematicians will not produce >2 "high SNR" papers per month as their rate calculated over any reasonable time scale; but their maximum per month? Very much the case that they will complete linked papers/dependencies/series of work/etc, spend a month in editing across multiple nearly-complete submissions, and so on.
And in mathematics, at least, priority on results is determined by the first time it appears out in the wild; for most, this is the arXiv.
Nobody is going to accept waiting a month on something that they fear may be scooped, costing them years of work and thought.
This affectively makes the arXiv obsolete for much of the purpose it has heretofore been put to.
AlotOfReading 7 hours ago [-]
Arxiv has had a limit of 3 simultaneous submissions for a couple years now. This is only adding per-month limits.
miki123211 5 hours ago [-]
I find it interesting that we, in the AI community, usually say "Arxiv", where mathematicians seem to say "the Arxiv".
Whereas I would say "Have they put this on Arxiv yet?", mathematicians would say "Have they put this on the Arxiv yet?".
hingler36 9 hours ago [-]
Do you have a concrete example of a high SNR researcher who will be limited by this policy? Aka one who is submitting more than two high quality papers per month?
sarjann 9 hours ago [-]
I imagine they could augment it with giving higher rates for some people based on things like h-index.
I hope something like this could also be adopted for some of our larger conferences - the absolute limits on co-authorship are what seem to cause the most grumbles.
> In September of 2016, arXiv received 9,869 submissions. In September of 2024, arXiv received 20,569 submissions. This September, arXiv received 40,363 submissions, which in turn generated almost 9,000 support tickets for arXiv staff and moderators.
> arXiv now limits submitters to up to two submissions per calendar month, with a limit of three total active submissions at any given time.
https://arxiv.org/abs/2601.13187
Another one of mine has been on hold for 18 days and counting.
I think it is a great service to the community, but clearly they are having a scalability problem.
The hosting cost side I can understand, though they received quite some funding recently, but I assume it's not going towards actually running the site, but to who knows what broader impacts and so on.
Moderation for Arxiv is a silly idea. It already shouldn't be taken as a quality signal that something is able to be up on Arxiv. Obviously they should remove illegal stuff, but beyond that, requiring moderation is a misunderstanding of their role and reason for popularity.
If hosting is too expensive, I guess an alternative aggregator could also arise. With just metadata and a hash of the pdf that can be hosted anywhere, and as the sumbitter, if you move the file, you can change the URL.
The main reason for Arxiv's existence is the timestamping and the easy referencing. (Though I admit that the stable hosting is also a pretty important part, but they mention moderation effort as the reason, not the hosting costs.)
Academia is losing sight of the forest for the trees, can't see more than an arm's length ahead of their noses.
It sounds like you're advocating for arXiv to become viXra. On https://vixra.org/all/2610 the second paper is currently "Correct Interpretation of the Great Discoveries in Particle Physics: I. Reconsidering the Higgs Boson via Vedic Vortex Structure", followed by a paper arguing that classical electrodynamics is bogus, "Gauss’s Flux Theorem Does not Hold in Time-Varying Electric Fields", and then "On the Boundary Problem in the Origins of Matter, Life, and Consciousness".
I don't think arXiv will continue to receive "quite some funding" if this is what it contains.
If Arxiv was the only game in town and it allowed everything, the result would be more like GitHub or Substack. That is, you wouldn't be able to trust everything hosted there, and a lot of it would be low-quality promotional garbage, with filtering moved to a different layer.
You could do some kind of karma system for example, where the prestige of your institution, your own karma, number of citations, papers published in high impact-factor journals etc would affect karma weight. You could have journals be like "awesome-x" lists on Github, with authors who have already been published there able to endorse other papers. I believe the mathematics community is trying that.
The story that academia likes to tell itself about itself is that it is skeptical and iconclastic and objectively evaluates each piece of science on its own virtues and it doesn't matter who says it. A wrong thing said by someone famous is still wrong, and truth spoken by a nobody is still truth. Of course reality is a bit further from this ideal. But outright admitting it would be hard.
It's a bit like the "in America, anyone can become president!".
But actually, thinking back, 10 years ago there may have been more of a stigma around a paper that was only on arxiv.
Not sure if you understand how the sausage gets made. People won't spend effort at scale (there are always exceptions) on stuff that's not incentivized, such as blog posts. If you work on something, the goal is to make it count. But if you fail on that, arXiv is a second best option.
Also, papers that get published at conferences also get out of date pretty fast anyway. That's just how fast moving fields work.
But things may end up moving the way you are saying, as long as incentives keep up.
Journals (1-2 years) -> conferences (6 months) -> arxiv (weeks to write once the work is done, then days to publish) -> next thing (immediate AI writeup, code, results, tests, open scripts, weights, full reproducibility, full access allowed for other AIs to scrape and learn from, and reorganize and recombine on a speed of days or hours from result to adoption)
One of weird aspects of CS research is the tendency to publish small papers. In more experimental fields, a paper may represent years of work for the first author. But CS students often finish multiple projects every year. Maybe it's because of the focus on conferences. You want to attend conferences (to build your networks, and because it's interesting and valuable), but you often don't get reimbursed if you don't have anything to present. Students from other fields can just present their work in progress, but CS students have to come up with publishable results.
Maybe CS should also strive towards bigger projects. Instead of publishing individual results (which someone often improves a year or two later), the papers could focus on the outcomes of entire research directions. That way you are more likely to produce something of lasting value. Something that is worth writing up and rewriting and polishing, until you understand the ideas much better than after the initial write-up.
By the same principle, someone in the early 2000s could reasonably say "only weirdos date online", but the mainstream can shift also.
Part of the reason that the value of Arxiv papers has gone up, is that it's pretty well known that the conference review system is already quite broken and random and lazy and superficial. So being rejected from there is not a good reason to ignore a paper, so conference-rejected Arxiv-only but good papers are a regular occurrence. If conference review had better signal, arxiv could have zero signal.
The other issue is the time delay. Conferences have stopped being an actual place to learn about new work, the way it had been 10+ years ago. Now it's all outdated stuff, and as a specialist you already know the important papers months before the actual conference, so the conference is a networking event basically. The real exchange of ideas moved to Arxiv and Github and social media, with its own problems, hype, algorithmic engagement optimization incentives etc.
But the pressure that shifted away from conferences to arxiv is continuing. The old guard was pearl clutching already by the erosion of conference rubber stamps as this big prestige. The newer ones are now trying to hold back the tide at the Arxiv line. But it will keep on moving faster and faster. And what is going to matter in the end is not how things used to be, but what delivers actual value. If the slop is slop, it will dwindle. If it starts to be actually good, all this will just turn into basically a dock-unions-opposing-automation story.
Arxiv papers are regularly cited and assumed to be suitable for peer review. So some standards need to be maintained . it's not a free for all.
I suspect a logical conclusion Arxiv and elsewhere may be an identity management system with an aggressive filter, and shared blacklists. I suspect that classifying people as spam/slop-submitters, then banning them (or whatever identity they used; name, email, name + organization etc), applying incremental rate limits over the general one, may be required.
Then, if someone posts rejected papers regularly and is deprioritized, it naturally follows that they may try to submit three more... and get stuck naturally. There is no need to ban them. They will still get a response, and they will get a fair shake - when there is time.
The situation right now is pretty crazy, for example this established TCS professor [1] has at least 4 arXiv submissions co-authored by him in September [2]. This includes the recent breakthrough on the matroid secretary problem, which has a pretty interesting AI story of its own, if you haven't seen it yet [3].
Of course prolific professors can work around this by working with younger researchers who upload the work, but this just makes the rate limit a solo author bottleneck, which feels a bit weird.
[1]: https://en.wikipedia.org/wiki/Mohammad_Hajiaghayi
[2]: https://arxiv.org/search/cs?query=Hajiaghayi&searchtype=auth...
[3]: Concurrent Discovery Disclosure: The proof of the main result in this manuscript was obtained in a conversation with ChatGPT-6 Astra on Tuesday, September 15, 2026 at 1:02 AM PDT. We then prepared this manuscript for public release, with the intent of uploading it on the morning of Thursday, September 17, 2026. In the early morning hours of September 17, while finalizing the submission, we discovered the manuscript of Abdi, Banihashem, Hajiaghayi, and Mittal, uploaded on September 16, 2026, which contains the same result via an essentially identical approach. We are sharing our manuscript nonetheless in case our exposition is of independent utility to the community, and we hope this experience stimulates broader discussion about concurrent discovery in the AI era. (From https://arxiv.org/abs/2609.20797 .)
But yes, it'd be nice if arxiv also rate-limited co-author submissions with some higher number to discourage the most flagrant PI co-authorship abuses.
And in mathematics, at least, priority on results is determined by the first time it appears out in the wild; for most, this is the arXiv.
Nobody is going to accept waiting a month on something that they fear may be scooped, costing them years of work and thought.
This affectively makes the arXiv obsolete for much of the purpose it has heretofore been put to.
Whereas I would say "Have they put this on Arxiv yet?", mathematicians would say "Have they put this on the Arxiv yet?".