Rendered at 19:57:57 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
dankai 1 days ago [-]
I've been working on research related to this in the context of diferent LLMs, and I can tell you that similarity != executability.
While you can make embeddings across different LLMs similar (i.e. universally looking) the few percentage of R2 that you are missing in translating the universal representation into a native representation are precisely those that make the hidden states executable in the LLM (which makes them useful).
They do this with simple embedding models because there the purpose of the embeddings is to measure similarity, but if you would like to use this principle to turn latent representations of one LLM into latent representations that are understandable/executably by a different LLM, you will fail.
I am not familiar with the standards of publishing in machine learning, but as someone trained in a mathematics background, this paper seems relatively light on details and heavy on exposition. Is that typical? Is this a really novel idea? Not trying to be snarky, just trying to understand how meaningful this is.
robrenaud 2 days ago [-]
It's cool that it proves that a bunch of vectorized outputs from an unknown embedder on an unknown dataset is in no way private, because of this ability to reverse engineer the embedder.
I talked to the author at his poster session at neurips and was able to get the gist, though I had read a lot about the platonic representation hypothesis, and this was one of my top 10 favorite papers in the conference.
1 days ago [-]
wging 2 days ago [-]
They link to their code on GitHub - see footnote 2 on page 2. I don't see it linked anywhere else, which makes it easy to miss. https://github.com/rjha18/vec2vec/
Nail2680 2 days ago [-]
It isn't a maths paper, so the conventions are different.
adastra22 2 days ago [-]
This isn’t math. The exposition IS the details.
rhelz 2 days ago [-]
You are not wrong. But this has by no means proven its up to the standard of being publishable in a machine learning journal. Its on arXiv.org, which, lets face it, at the end of the day is a vanity press.
contubernio 2 days ago [-]
Calling the arxiv a vanity press shows you aren't a researcher. In math and physics all the best stuff is on the arxiv and the general level is well above the level of most journals. Journals mainly serve as accreditation and many are basically mediocre - the review process resta more value than it adds overall.
rhelz 1 days ago [-]
What is your definition of a vanity press? Mine is that it will publish anything from anybody.
Sure, there's lots of good stuff on there. But there's lots which is not good. Publishing there is not doing science. Science requires peer review.
Its a wonderful resource, but when overworked journalists (or worse, AI) just picks up a paper from there and writes a breathless article on how cool this new idea is, it is misleading to the general public.
Worse, it can get the general public (and from the comments here, even people working in the field) to forget what science is and why arXiv is not a scientific journal.
canjobear 2 days ago [-]
It was accepted to NeurIPS.
efavdb 2 days ago [-]
At a minimum posting to arxiv gives others a standard way to cite the work.
odyssey7 2 days ago [-]
The pace of things is moving along so rapidly right now, I’m not sure that waiting for peer reviews is always a wise move. Doubly so if there’s a paywall; why limit your article’s impact by placing it where practitioners’ agents might not be able to access it? The rapid progress right now is challenging for conventional academic processes.
If the value of the paper is difficult to independently verify, for example, if it depends on the credibility of the author, then the academic ritual can add something. If it’s a mathematical result, one that can be automatically verified, or a machine learning technique that anyone can try with Claude code reconstructing it for them, this sort of pre-print publishing model is advantageous.
rhelz 2 days ago [-]
// why limit your article's impact //
Because...science? It's not science until it passes peer review.
I'm not advocating that everybody stops posting to arXiv, and I'm not saying you can't find good stuff there. I'm just saying, it's a vanity press, there is absolutely no guarantee of the paper's quality.
And being published by a famous professor from a prestigious university is also no guarantee. If we've learned anything from the non-reproducibility crisis, it is that a paper's origin story is no guarantee.
Xmd5a 1 days ago [-]
> It's not science until it passes peer review.
You mean Robert Maxwell's quasi-monopoly on scientific publications ?
Then again neither is peer review. Reproducing research is the only way to prove reproducibility and thereby lend credibility to the claims.
raattgift 1 days ago [-]
> It's not science until it passes peer review
Peer reviewers don't generally reproduce the work in the publications they are asked to review, at least not during the review process itself. Neither do journal editors.
If you want a real vanity press, check vixra and its origin story.
ArXiv does have moderators and endorsers in each area, although they are usually light-touch, weeding out literally unreadable submissions, ones that so obviously ignore formatting guidelines that it beggars belief they could comply with a typical journal's rules, and ones that are clearly submitted to the wrong area.
Having a stable document early (a pre-print) tends to widen the scrutiny of papers that may ultimately be published to a journal whose editor's expertise lies in a very different area from the submission; this can and does lead to author corrections being made before journal publication.
There are plenty of peer-reviewed papers which are hot garbage that have found their way into prestigious high-impact journals like Nature and Science (see <https://retractionwatch.com/the-retraction-watch-leaderboard...> for examples, and note none of the top 10 are in areas covered by the arXiv).
> being published by a famous professor from a prestigious university is also no guarantee
Everyone in academia knows this, including the vast majority of famous professors from prestigious universities, because of online repositories like <https://retractionwatch.com/retractions-by-nobel-prize-winne...> (and much more sadly because of <https://en.wikipedia.org/wiki/Nobel_disease>, lower-profile versions of which academic paper-writers -- and dissertation writers -- tend to encounter as they chase the history of the problem before them or read late citations to works they are relying upon). This particular part of your set of claims is is not a real problem in academia or with the arXiv in particular.
Finally, what value does publication in a predatory journal bring? Do you believe that peer review and editing were actually even performed in the majority of MDPI's most predatory journals, for example? <https://www.predatoryjournals.org/news/list-of-all-mdpi-pred...> Some papers published in some of their pay-for-publication open access journals don't even get submitted to the arXiv; one might hope this is out of embarrassment by the authors, although the (low) bar set by the arXiv itself is certainly a factor.
The remedy for a reasonably argued but wrong academic paper isn't lack of publication, failed peer review, or editorial alteration, but rather reply papers. That's the academic dialogue.
rhelz 1 days ago [-]
Thanks for this reply, your points are well taken and it really got me to think about this stuff.
I suppose you could view arxiv as being a way to let more people do a peer review on the papers, i.e. as one component of a more open and transparent peer review process. As such, it is certainly a welcome alternative to what the snootier journal reviews do, which is to put off reading your paper for a year, and then spend about 50 seconds looking for a reason to say "no".
But I had a narrow point to make when I said that arxiv is a vanity press: just because something is on there doesn't guarantee that it is correct, or even useful. It is a vanity press in the sense that it will publish pretty much anything--modulo obscenity and porn laws, and apparently also some sanity checking by arxiv moderators--but so does a vanity press for print books.
To your point of predatory journals, power-broker reviewers and such, yes, I have done my fair share of suffering from them. It slows down good research and passes through bad research. The system is badly in need of reform, and arxiv gives a much needed alternative.
But really, the problem is (as one of my profs said) that science is totally run on volunteer, unpaid labor. Peer reviewers are very busy, get paid nothing, it is mostly a distraction. This leads them to look for other ways of being compensated--e.g. the can try to be a power broker, a feared authority, etc.
The real solution would be to actually pay peer reviewers and fund them so they can actually reproduce the results. Here's how it could work: when a researcher submits a grant proposal, they budget for enough money to fund the research and enough money to peer review and reproduce it. If the research isn't good enough to warrant an attempt to validate it, then it isn't worth doing in the first place.
raattgift 4 hours ago [-]
Also
> it really got me to think about this stuff.
Thank you, that is very generous of you to say.
rhelz 1 hours ago [-]
Thanks. Nobody ever admits when they are wrong on the internet, so I'm trying to be the change I want to see.
W.R.T. the grand historical tradition, etc, note that the great majority of people who practiced that were, like Plato, independently wealthy and otherwise idle. Those who didn't need any money.
That's why we give Judges, teachers, and Professors tenure, a guaranteed job for life. People are more likely to be fair judges of other people's performance when they don't need the money.
So I guess if you want a non-conservative fix, it would be universal guaranteed minimum income. Something we already have, BTW, but you have to be over 62 to get it.
We've tried to open it up to more than just the 1%, but we didn't really go all the way. It is really unreasonable to expect somebody who isn't independently wealthy to do free work. It is a lot of work to properly review a paper, and reproducing the results is something you simply can't do at all if you are supposed to already be working 70 hours a week generating your own research.
Of course a researcher paying reviewers would be a disastrous idea. But if the grant-awarding body would hire 3rd parties to validate that their money was actually well spent, I think it would be a win/win.
Thanks for the references on where paying reviewers has already been tried. I'd like to see others reproduce their success :-)
raattgift 4 hours ago [-]
No paper is guaranteed to be completely accurate and correct at the time of publication; the remedy for honest errors is follow-on papers by the original authors or others in the field putting their names to formalized counter-arguments.
This is how the academic dialogue works, and it dates back to early durable records in the ancient world -- about as far before Plato as Plato's Academy is before today's universities. Innovations like academic journals come from "merely" four centuries ago, and pre-publication evaluation by peers about three hundred years ago. Anonymized pre-publication review is more recent still, and publishing schedules (as well as the needs of the submitter to have the submitted work dealt with in good time) don't admit a dialogue over some subtle point a referee takes some issue with. Moreover, such a pre-publication dialogue is generally non-public, and often also filtered through an editor. It can get messy [1].
As I'm sure you've discovered in your own post-secondary education, academia is very conservative (in the sense of "retaining what mostly works, even in the face of mounting scaling problems", foremost among those being the ever-increasing volume of publishable works generated by researchers).
Direct compensation of referees by journals has a number of pitfalls ranging from the fiddly details of how to deal with income tax liability pay (and credits towards future publication fees, etc) incurs to whether it compromises the objectivity of pre-publication evaluation (which is frankly already not very good in highly specialized areas).
Payment of reviewers by paper submitters is imho an even worse idea -- how do you preserve the anonymity of referees (especially with respect to the submitters)? Do submitters or reviewers have exposure to their (probably different) tax authorities for these transactions? (I can think of at least four national tax agencies that would view the received pay as taxable income, and also potentially subject to value-added tax on professional services. One already sees this problem with conference honoraria.)
Another -- one that already exists in the real world -- might be for institutions and grant-funders to build in support for peer review of others' work by project members. To the extent that detailed pre-publication review benefits academia, and societies that fund academia, as a whole, maybe referees should be rewarded (or even obliged) to do paid reviewing duty. One might compare the requirement to do pro bono work imposed on many practising lawyers by their regulating authorities (especially interesting are jurisdictions in which a party losing a court contest to a party represented pro bono usually must pay into a fund that supports the overall scheme).
Personally I think it would be more useful to everyone to have a solid reply paper published in due course than a blocked initial publication. My ideal is that referees aim to help submitters not accidentally professionally embarrass themselves.
Then on your reproduction idée fixe, supporting the generation of papers which confirm reproduction is something that societies could certainly work on. Many incentives to write and submit papers calling out flaws in others' work already exist, and of course could be expanded. I don't think any of that should be done with masking of the participants. Additionally, reproduction of results is often less interesting than (sometimes much) later testing of published results in significantly different ways.
(Indeed, foundational papers from more than ninety years ago are often re-proven by modern-day experiments, both as a side effect of (or opportunity arising from) pursuing the primary research goal, and as a way of testing e.g. new metrology technology, new computational tools, mathematical advances, and so on. Also, some of those foundational papers simply could not be tested with great accuracy with tools existing at the time of publication. Their publication, however, tended to drive the development of such tools, and in modern days such foundational papers are found to be good only within certain limits that could not have been explored closer to the time of publication.)
One way to pose/(think about) the problem is that there are two finite metric spaces linked by an unknown odometry (damn you autocorrect). The problem is to recover that unknown isometry.
This, like graph isometry, can be very computationally intensive in the worst case. However, heuristics to aid matching one vertex on one graph to another vertex on another graph using local, semilocal structural signatures can be very effective on particular cases.
One can of course argue that the spaces are not designed as metric spaces. Even if true, these might be metrizable topological spaces.
More generally, if these are indeed non-metric spaces one can still pose it as finding the unknown isomorphism between two poset spaces.
In my other comment I was using the property of maximal chains -- Identify the longest chains in both posets. The isomorphism must map the longest chain in Poset 1 directly to a longest chain in Poset 2, preserving the exact linear order.
srean 2 days ago [-]
Let's assume that monotonocity of pair-wise distances are preserved.
Without knowing the details of how the paper solved the problem, my first attempt would be to find the diametrically distant pair of points in the two different embeddings and assume that the pair is the same pair. Then find the next distant pairs and so on.
After sufficiently many such pairs have been found, or better still, the largest d-simplex is found, find that scaled rigid body transformation that makes the corresponding pairs coincide. Proceeding this way ought to be less work than solving a generic graph isomorphism problem.
robrenaud 2 days ago [-]
I think a less stringent, but still workable assumption is that for very similair objects, their distances will be small. This is much easier to accomplish than agreement across all pairs.
srean 1 days ago [-]
Could you explain a bit more. What you say about similar objects is obviously true. However the algorithm sketch that you have in your mind is a little implicit. Could you make it more explicit. I am quite curious.
I'm not as heavy on the maths stuff involved in this as other people commenting appear to be.
But the idea makes sense, of course there is still recoverable data in embeddings, that's the point. Though as I constantly find the more you try to squeeze into an n bit vector the more watered down everything gets.
I suppose a latent space could be encrypted/mapped in some way to resolve that, but how many people are exposing their vectors in the first place?
ViscountPenguin 1 days ago [-]
The point of the paper isn't that embeddings contain information, it's that even if you don't know what model generated a set of embeddings you can still recover information from the geometry of the point cloud itself.
The fact that this is possible also adds some pretty strong restriction to the set of possible maps you could use to remove that information. No linear map will work since all embedding spaces are ~an orthonormal matrix apart, so some form of encryption is necessary. This wasn't known until very recently.
fennecfoxy 6 hours ago [-]
Ah right, thank you for clarifying.
But I think my point still stands, isn't the geometry information THE information I referred to in the first place? Obviously the vector size gives you the granularity but it's kind of unavoidable to positionally encode information in a latent space...that's literally what they're for?
But yes, it is very cool to know that regardless of exact implementation finding x,y,z representations of some dataset with various relationships (like language) creates similar geometry/clues across all the implementations.
stephantul 2 days ago [-]
I’ve never liked that this was called “the platonic representation hypothesis”.
Lots of weird baggage attached and seems like a waste of a good name.
pksebben 1 days ago [-]
"We believe these representations are not serious, they're just really good friends."
FloorEgg 1 days ago [-]
This is being positioned as evidence of the Platonic Representation Hypothesis, but isn't it more likely an artifact of training models on the same/similar data sets?
rhelz 2 days ago [-]
Cyberphrenology. In any two random graphs, you'll find an isomorphic graph which is can be up to log of the size of the graphs.
And if the LLM has been trained up to the limit of what data it can hold, it is going to be random. Proof below if it isn't obvious.
The entire effort of all people who are trying to understand how LLMs work, how they represent their data, its all bound to fail.
Proof: a LLM is a very good approximation of the Solomonov/Levin/Kolmogorov universal probability function on tokens. As such, it will be random--pure white noise--because if you found any patterns in there, you could exploit the regularity and come up with a smaller set of weights for the same LLM.
There are no patterns there to be found. They have all been factored out by training the neural net until it couldn't learn any more.
sdenton4 2 days ago [-]
/a smaller set of weights for the same LLM./
Distillation is alive and well... Earlier work on model printing also found that it's pretty easy to find smaller sets of parameters which can replicate the behavior of the entire network with pretty good fidelity.
Large parameter counts give space to explore, and give routes out of what would be local minima in a lower dimensional space.
In other words, there's no guarantee that any given trained model is a minimal representation of its training set.
rhelz 2 days ago [-]
I'm not claiming any arbitrary set of weights is a minimal representation. But typically, if people could achieve the same quality of results with a smaller set of weights, or weights which have been quantized to lower bit representations, etc, they would have published the smaller one instead.
andrewflnr 2 days ago [-]
You kind of are claiming they're minimal, though. Because if they're not, your statement that "if you found any patterns in there, you could exploit the regularity..." implies nothing. Yeah, the patterns are there, and people are exploiting them.
Your socioeconomic argument just doesn't hold either. People don't delay releasing models until they've minimized it to the theoretical limit. They ship it when it's good enough for whatever job they're making it for.
rhelz 1 days ago [-]
If somebody finds patterns in them, that means they can be improved by reducing the patterns. Patterns and symmetry are forms of redundancy.
The less memory the weights take to meet a level of competency, the more random, and therefore patternless and inscrutable they will be.
andrewflnr 1 days ago [-]
No kidding. Clearly there's some redundancy, or however you want to frame it, in real-world released models. That still means your argument for why they can never be interpretable is based an incorrect premise.
canjobear 2 days ago [-]
The weights aren’t compressed. So there are interpretable redundancies in practice.
rhelz 2 days ago [-]
If the weights arn't compressed, then a smaller set of weights would perform as well. Sure, you can always induce as much symmetry and patterns as you want by bloating the data set, but that hardly gives us insight into how a set of weights which is "as full as it can be" of information.
canjobear 2 days ago [-]
The point of TFA is that there are regularities you can exploit in the actually existing weights of machine learning systems, not in some hypothetically maximally efficient weights. The maximally efficient weights would indeed have no structure, but that’s not what anyone is working with.
measurablefunc 2 days ago [-]
What is the (co)homology of this space?
chombier 2 days ago [-]
That of the underlying, hypothetical universal brain topology?
While you can make embeddings across different LLMs similar (i.e. universally looking) the few percentage of R2 that you are missing in translating the universal representation into a native representation are precisely those that make the hidden states executable in the LLM (which makes them useful).
They do this with simple embedding models because there the purpose of the embeddings is to measure similarity, but if you would like to use this principle to turn latent representations of one LLM into latent representations that are understandable/executably by a different LLM, you will fail.
Note this is version 4 of the paper and the original post was version 1 (I think?)
OpenReview (for NeurIPS) for the curious: https://openreview.net/forum?id=jiCLUPq5xv
I talked to the author at his poster session at neurips and was able to get the gist, though I had read a lot about the platonic representation hypothesis, and this was one of my top 10 favorite papers in the conference.
Sure, there's lots of good stuff on there. But there's lots which is not good. Publishing there is not doing science. Science requires peer review.
Its a wonderful resource, but when overworked journalists (or worse, AI) just picks up a paper from there and writes a breathless article on how cool this new idea is, it is misleading to the general public.
Worse, it can get the general public (and from the comments here, even people working in the field) to forget what science is and why arXiv is not a scientific journal.
If the value of the paper is difficult to independently verify, for example, if it depends on the credibility of the author, then the academic ritual can add something. If it’s a mathematical result, one that can be automatically verified, or a machine learning technique that anyone can try with Claude code reconstructing it for them, this sort of pre-print publishing model is advantageous.
Because...science? It's not science until it passes peer review.
I'm not advocating that everybody stops posting to arXiv, and I'm not saying you can't find good stuff there. I'm just saying, it's a vanity press, there is absolutely no guarantee of the paper's quality.
And being published by a famous professor from a prestigious university is also no guarantee. If we've learned anything from the non-reproducibility crisis, it is that a paper's origin story is no guarantee.
You mean Robert Maxwell's quasi-monopoly on scientific publications ?
https://www.theguardian.com/science/2017/jun/27/profitable-b...
Peer reviewers don't generally reproduce the work in the publications they are asked to review, at least not during the review process itself. Neither do journal editors.
If you want a real vanity press, check vixra and its origin story.
ArXiv does have moderators and endorsers in each area, although they are usually light-touch, weeding out literally unreadable submissions, ones that so obviously ignore formatting guidelines that it beggars belief they could comply with a typical journal's rules, and ones that are clearly submitted to the wrong area.
If anything, I think sometimes the net is too fine. Will Kinney just this week had a version of <https://www.acsu.buffalo.edu/~whkinney/SpecialRelativityBoot...> rejected by the arXiv, for example. The reasons for arXiv-rejection can be opaque.
Having a stable document early (a pre-print) tends to widen the scrutiny of papers that may ultimately be published to a journal whose editor's expertise lies in a very different area from the submission; this can and does lead to author corrections being made before journal publication.
There are plenty of peer-reviewed papers which are hot garbage that have found their way into prestigious high-impact journals like Nature and Science (see <https://retractionwatch.com/the-retraction-watch-leaderboard...> for examples, and note none of the top 10 are in areas covered by the arXiv).
> being published by a famous professor from a prestigious university is also no guarantee
Everyone in academia knows this, including the vast majority of famous professors from prestigious universities, because of online repositories like <https://retractionwatch.com/retractions-by-nobel-prize-winne...> (and much more sadly because of <https://en.wikipedia.org/wiki/Nobel_disease>, lower-profile versions of which academic paper-writers -- and dissertation writers -- tend to encounter as they chase the history of the problem before them or read late citations to works they are relying upon). This particular part of your set of claims is is not a real problem in academia or with the arXiv in particular.
Finally, what value does publication in a predatory journal bring? Do you believe that peer review and editing were actually even performed in the majority of MDPI's most predatory journals, for example? <https://www.predatoryjournals.org/news/list-of-all-mdpi-pred...> Some papers published in some of their pay-for-publication open access journals don't even get submitted to the arXiv; one might hope this is out of embarrassment by the authors, although the (low) bar set by the arXiv itself is certainly a factor.
The remedy for a reasonably argued but wrong academic paper isn't lack of publication, failed peer review, or editorial alteration, but rather reply papers. That's the academic dialogue.
I suppose you could view arxiv as being a way to let more people do a peer review on the papers, i.e. as one component of a more open and transparent peer review process. As such, it is certainly a welcome alternative to what the snootier journal reviews do, which is to put off reading your paper for a year, and then spend about 50 seconds looking for a reason to say "no".
But I had a narrow point to make when I said that arxiv is a vanity press: just because something is on there doesn't guarantee that it is correct, or even useful. It is a vanity press in the sense that it will publish pretty much anything--modulo obscenity and porn laws, and apparently also some sanity checking by arxiv moderators--but so does a vanity press for print books.
To your point of predatory journals, power-broker reviewers and such, yes, I have done my fair share of suffering from them. It slows down good research and passes through bad research. The system is badly in need of reform, and arxiv gives a much needed alternative.
But really, the problem is (as one of my profs said) that science is totally run on volunteer, unpaid labor. Peer reviewers are very busy, get paid nothing, it is mostly a distraction. This leads them to look for other ways of being compensated--e.g. the can try to be a power broker, a feared authority, etc.
The real solution would be to actually pay peer reviewers and fund them so they can actually reproduce the results. Here's how it could work: when a researcher submits a grant proposal, they budget for enough money to fund the research and enough money to peer review and reproduce it. If the research isn't good enough to warrant an attempt to validate it, then it isn't worth doing in the first place.
> it really got me to think about this stuff.
Thank you, that is very generous of you to say.
W.R.T. the grand historical tradition, etc, note that the great majority of people who practiced that were, like Plato, independently wealthy and otherwise idle. Those who didn't need any money.
That's why we give Judges, teachers, and Professors tenure, a guaranteed job for life. People are more likely to be fair judges of other people's performance when they don't need the money.
So I guess if you want a non-conservative fix, it would be universal guaranteed minimum income. Something we already have, BTW, but you have to be over 62 to get it.
We've tried to open it up to more than just the 1%, but we didn't really go all the way. It is really unreasonable to expect somebody who isn't independently wealthy to do free work. It is a lot of work to properly review a paper, and reproducing the results is something you simply can't do at all if you are supposed to already be working 70 hours a week generating your own research.
Of course a researcher paying reviewers would be a disastrous idea. But if the grant-awarding body would hire 3rd parties to validate that their money was actually well spent, I think it would be a win/win.
Thanks for the references on where paying reviewers has already been tried. I'd like to see others reproduce their success :-)
This is how the academic dialogue works, and it dates back to early durable records in the ancient world -- about as far before Plato as Plato's Academy is before today's universities. Innovations like academic journals come from "merely" four centuries ago, and pre-publication evaluation by peers about three hundred years ago. Anonymized pre-publication review is more recent still, and publishing schedules (as well as the needs of the submitter to have the submitted work dealt with in good time) don't admit a dialogue over some subtle point a referee takes some issue with. Moreover, such a pre-publication dialogue is generally non-public, and often also filtered through an editor. It can get messy [1].
As I'm sure you've discovered in your own post-secondary education, academia is very conservative (in the sense of "retaining what mostly works, even in the face of mounting scaling problems", foremost among those being the ever-increasing volume of publishable works generated by researchers).
Your own idea is quite conservative, and has been tried from time to time. Direct compensation by journals of referees for their time has been the subject of rigorous human experimentation: https://10.1097/CCM.0000000000006637 https://www.biorxiv.org/content/10.1101/2025.03.18.644032v1 as recent examples, with mixed results.
Direct compensation of referees by journals has a number of pitfalls ranging from the fiddly details of how to deal with income tax liability pay (and credits towards future publication fees, etc) incurs to whether it compromises the objectivity of pre-publication evaluation (which is frankly already not very good in highly specialized areas).
Payment of reviewers by paper submitters is imho an even worse idea -- how do you preserve the anonymity of referees (especially with respect to the submitters)? Do submitters or reviewers have exposure to their (probably different) tax authorities for these transactions? (I can think of at least four national tax agencies that would view the received pay as taxable income, and also potentially subject to value-added tax on professional services. One already sees this problem with conference honoraria.)
Another -- one that already exists in the real world -- might be for institutions and grant-funders to build in support for peer review of others' work by project members. To the extent that detailed pre-publication review benefits academia, and societies that fund academia, as a whole, maybe referees should be rewarded (or even obliged) to do paid reviewing duty. One might compare the requirement to do pro bono work imposed on many practising lawyers by their regulating authorities (especially interesting are jurisdictions in which a party losing a court contest to a party represented pro bono usually must pay into a fund that supports the overall scheme).
Personally I think it would be more useful to everyone to have a solid reply paper published in due course than a blocked initial publication. My ideal is that referees aim to help submitters not accidentally professionally embarrass themselves.
Then on your reproduction idée fixe, supporting the generation of papers which confirm reproduction is something that societies could certainly work on. Many incentives to write and submit papers calling out flaws in others' work already exist, and of course could be expanded. I don't think any of that should be done with masking of the participants. Additionally, reproduction of results is often less interesting than (sometimes much) later testing of published results in significantly different ways.
(Indeed, foundational papers from more than ninety years ago are often re-proven by modern-day experiments, both as a side effect of (or opportunity arising from) pursuing the primary research goal, and as a way of testing e.g. new metrology technology, new computational tools, mathematical advances, and so on. Also, some of those foundational papers simply could not be tested with great accuracy with tools existing at the time of publication. Their publication, however, tended to drive the development of such tools, and in modern days such foundational papers are found to be good only within certain limits that could not have been explored closer to the time of publication.)
- --
[1] https://theconversation.com/hate-the-peer-review-process-ein... (2014)
This, like graph isometry, can be very computationally intensive in the worst case. However, heuristics to aid matching one vertex on one graph to another vertex on another graph using local, semilocal structural signatures can be very effective on particular cases.
One can of course argue that the spaces are not designed as metric spaces. Even if true, these might be metrizable topological spaces.
More generally, if these are indeed non-metric spaces one can still pose it as finding the unknown isomorphism between two poset spaces.
In my other comment I was using the property of maximal chains -- Identify the longest chains in both posets. The isomorphism must map the longest chain in Poset 1 directly to a longest chain in Poset 2, preserving the exact linear order.
Without knowing the details of how the paper solved the problem, my first attempt would be to find the diametrically distant pair of points in the two different embeddings and assume that the pair is the same pair. Then find the next distant pairs and so on.
After sufficiently many such pairs have been found, or better still, the largest d-simplex is found, find that scaled rigid body transformation that makes the corresponding pairs coincide. Proceeding this way ought to be less work than solving a generic graph isomorphism problem.
I explained my thoughts in a comment here
https://news.ycombinator.com/item?id=49595424
But the idea makes sense, of course there is still recoverable data in embeddings, that's the point. Though as I constantly find the more you try to squeeze into an n bit vector the more watered down everything gets.
I suppose a latent space could be encrypted/mapped in some way to resolve that, but how many people are exposing their vectors in the first place?
The fact that this is possible also adds some pretty strong restriction to the set of possible maps you could use to remove that information. No linear map will work since all embedding spaces are ~an orthonormal matrix apart, so some form of encryption is necessary. This wasn't known until very recently.
But I think my point still stands, isn't the geometry information THE information I referred to in the first place? Obviously the vector size gives you the granularity but it's kind of unavoidable to positionally encode information in a latent space...that's literally what they're for?
But yes, it is very cool to know that regardless of exact implementation finding x,y,z representations of some dataset with various relationships (like language) creates similar geometry/clues across all the implementations.
And if the LLM has been trained up to the limit of what data it can hold, it is going to be random. Proof below if it isn't obvious.
The entire effort of all people who are trying to understand how LLMs work, how they represent their data, its all bound to fail.
Proof: a LLM is a very good approximation of the Solomonov/Levin/Kolmogorov universal probability function on tokens. As such, it will be random--pure white noise--because if you found any patterns in there, you could exploit the regularity and come up with a smaller set of weights for the same LLM.
There are no patterns there to be found. They have all been factored out by training the neural net until it couldn't learn any more.
Distillation is alive and well... Earlier work on model printing also found that it's pretty easy to find smaller sets of parameters which can replicate the behavior of the entire network with pretty good fidelity.
Large parameter counts give space to explore, and give routes out of what would be local minima in a lower dimensional space.
In other words, there's no guarantee that any given trained model is a minimal representation of its training set.
Your socioeconomic argument just doesn't hold either. People don't delay releasing models until they've minimized it to the theoretical limit. They ship it when it's good enough for whatever job they're making it for.
The less memory the weights take to meet a level of competency, the more random, and therefore patternless and inscrutable they will be.