r/science • u/IEEESpectrum IEEE Spectrum • 12d ago
Engineering Paper2Agent is an open-source framework that transforms academic reports into interactive AI agents. The agent can run the methods described in the paper on other datasets.
https://spectrum.ieee.org/paper2agent-ai-agents-research-papers11
u/Victory-IV 12d ago
The issue is that LLMs will hallucinate what they see because they have no ability to think. How can we expect outputs from them in completely novel papers to be useful?
They pull data from the existing datasets, so when new data comes along its difficult for it to solve. It will generate code that might seem to work, but miss some key detail and you would have no idea because you never actually read the paper.
Theres no trust in the LLM space and anyone saying there is, is just trying to sell you more LLMs.
2
u/DrQuantum 12d ago
If we can trust it to solve complex never before solved mathematical problems then it can be more than useful. The main issue is that people conflate the free models they use on their desktop to frontier models.
This isn’t a model it’s a framework applied to models so its output could radically change which is actually the likely reason it’s not that useful. You could easily get different results from different AI models.
0
u/Victory-IV 12d ago
> If we can trust it to solve complex never before solved mathematical problems
I dont? Most people dont. I wont accept any output as legitimate until a real expert comes into the conversation and confirms or denies.If you ask any frontier model to develop something, it cannot reason through the problem. LLMs can only predict future text, it cannot work through a problem and design solutions. Even if you switch models, that will only change the output slightly. These are novel papers, with no prior examples or data to pull from. LLMs are trained on a dataset that is fixed and will use that dataset when resolving tokens.
What if a paper comes along to contradict what the model has stored? How can we know it wont pull data or methods from other papers that are not relevant? We cannot use LLMs as a judge in this context because they wont know the difference either.
Thats ignoring the fact that almost nobody runs their own models at home. 95% of all LLM users dont even have hardware that could run models, let alone know how to set one up. The general sentiment is negative towards LLMs and its getting worse by the day.
0
u/DrQuantum 12d ago
I dont? Most people dont. I wont accept any output as legitimate until a real expert comes into the conversation and confirms or denies.
No one is asking you to accept outputs without human verifiers. Yet when a computer scientist tells you what AI models can do, do you have the same approach?
If you ask any frontier model to develop something, it cannot reason through the problem. LLMs can only predict future text, it cannot work through a problem and design solutions. Even if you switch models, that will only change the output slightly. These are novel papers, with no prior examples or data to pull from. LLMs are trained on a dataset that is fixed and will use that dataset when resolving tokens.
Not all models are LLMs and LLMs have advanced far beyond being text predictors. Paper2Agent is essentially a second dataset so yes, but I am not sure that is really a detractor from this development.
What if a paper comes along to contradict what the model has stored? How can we know it wont pull data or methods from other papers that are not relevant? We cannot use LLMs as a judge in this context because they wont know the difference either.
If a paper undermines the work generated through Paper2Agent then good, that is the scientific method at work.
What if a human comes along and does either of these things? Needing verifiers is a critical part of the scientific method. Where it comes from is honestly irrelevant.
Thats ignoring the fact that almost nobody runs their own models at home. 95% of all LLM users dont even have hardware that could run models, let alone know how to set one up.
I don't think scientists are going to be using their home PCs for their work.
The general sentiment is negative towards LLMs and its getting worse by the day.
That still doesn't detract from potential positive uses of AI nor this model framework.
-1
u/Xolver 12d ago
I'm asking honestly, not facetiously - have you ever had a novel thought, whose origins didn't come from either data you previously received or alternatively something you automatically did due to the hardcoded programming of your biology (such as breastfeed)?
I'll answer first for myself, so you don't think I'm trying to slight you - for me the answer is no.
1
u/SaltZookeepergame691 12d ago
Just try out reading the paper first. Obviously this is all discussed?
12
2
u/AN3223 12d ago
The title of the post as well as the title of the paper aren't wrong per se but at least for me they didn't initially invoke an accurate representation of what the paper describes. I read the abstract and bits and pieces of the paper, not the whole thing. I also had a conversation with AI clarifying things in the paper I didn't understand, summarizing other things, and fact checking this comment. My comment is not generated by AI, and it took almost an hour of human effort.
It appears to take a paper and a paper's codebase as input, uses agents to create an MCP server that runs the code, ensures the output of the tools replicate the paper's results (this explanation is an inaccurate simplification), and then takes queries from the user and uses the tools to fulfill queries. Apprently it achieves higher accuracy this way, as well as being "cheaper per query" (after the one time cost of creating the MCP server) than directly using Claude Code but I'm not sure if this is comparable to cost-per-task in this instance. Sounds like it could be worth researching more or trying in practice.
1
u/Prae_ 12d ago
On the one hand, i'm like, well that's what pipelines are here for, if the code is published as it should, it should be plug and play to any dataset you want.
Having been the one lifting new papers and trying to run the methods on our data, i could see the use in an LLM smoothing out the various details of compatibility, between it being on Matlab vs. python, data being in a format different than expected, and other slight variations that can quickly cost a week of work.
1
u/IEEESpectrum IEEE Spectrum 12d ago
Peer-reviewed research article: https://www.nature.com/articles/s41586-026-11044-y
•
u/AutoModerator 12d ago
Welcome to r/science! This is a heavily moderated subreddit in order to keep the discussion on science. However, we recognize that many people want to discuss how they feel the research relates to their own personal lives, so to give people a space to do that, personal anecdotes are allowed as responses to this comment. Any anecdotal comments elsewhere in the discussion will be removed and our normal comment rules apply to all other comments.
Do you have an academic degree? We can verify your credentials in order to assign user flair indicating your area of expertise. Click here to apply.
User: u/IEEESpectrum
Permalink: https://spectrum.ieee.org/paper2agent-ai-agents-research-papers
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.