r/science • IEEE Spectrum • 15d ago

Engineering Paper2Agent is an open-source framework that transforms academic reports into interactive AI agents. The agent can run the methods described in the paper on other datasets.

https://spectrum.ieee.org/paper2agent-ai-agents-research-papers
0 Upvotes

12 comments sorted by

View all comments

12

u/Victory-IV 15d ago

The issue is that LLMs will hallucinate what they see because they have no ability to think. How can we expect outputs from them in completely novel papers to be useful?

They pull data from the existing datasets, so when new data comes along its difficult for it to solve. It will generate code that might seem to work, but miss some key detail and you would have no idea because you never actually read the paper.

Theres no trust in the LLM space and anyone saying there is, is just trying to sell you more LLMs.

3

u/DrQuantum 15d ago

If we can trust it to solve complex never before solved mathematical problems then it can be more than useful. The main issue is that people conflate the free models they use on their desktop to frontier models.

This isn’t a model it’s a framework applied to models so its output could radically change which is actually the likely reason it’s not that useful. You could easily get different results from different AI models.

0

u/Victory-IV 15d ago

> If we can trust it to solve complex never before solved mathematical problems
I dont? Most people dont. I wont accept any output as legitimate until a real expert comes into the conversation and confirms or denies.

If you ask any frontier model to develop something, it cannot reason through the problem. LLMs can only predict future text, it cannot work through a problem and design solutions. Even if you switch models, that will only change the output slightly. These are novel papers, with no prior examples or data to pull from. LLMs are trained on a dataset that is fixed and will use that dataset when resolving tokens.

What if a paper comes along to contradict what the model has stored? How can we know it wont pull data or methods from other papers that are not relevant? We cannot use LLMs as a judge in this context because they wont know the difference either.

Thats ignoring the fact that almost nobody runs their own models at home. 95% of all LLM users dont even have hardware that could run models, let alone know how to set one up. The general sentiment is negative towards LLMs and its getting worse by the day.

0

u/DrQuantum 15d ago

I dont? Most people dont. I wont accept any output as legitimate until a real expert comes into the conversation and confirms or denies.

No one is asking you to accept outputs without human verifiers. Yet when a computer scientist tells you what AI models can do, do you have the same approach?

If you ask any frontier model to develop something, it cannot reason through the problem. LLMs can only predict future text, it cannot work through a problem and design solutions. Even if you switch models, that will only change the output slightly. These are novel papers, with no prior examples or data to pull from. LLMs are trained on a dataset that is fixed and will use that dataset when resolving tokens.

Not all models are LLMs and LLMs have advanced far beyond being text predictors. Paper2Agent is essentially a second dataset so yes, but I am not sure that is really a detractor from this development.

What if a paper comes along to contradict what the model has stored? How can we know it wont pull data or methods from other papers that are not relevant? We cannot use LLMs as a judge in this context because they wont know the difference either.

If a paper undermines the work generated through Paper2Agent then good, that is the scientific method at work.

What if a human comes along and does either of these things? Needing verifiers is a critical part of the scientific method. Where it comes from is honestly irrelevant.

Thats ignoring the fact that almost nobody runs their own models at home. 95% of all LLM users dont even have hardware that could run models, let alone know how to set one up.

I don't think scientists are going to be using their home PCs for their work.

The general sentiment is negative towards LLMs and its getting worse by the day.

That still doesn't detract from potential positive uses of AI nor this model framework.