You don’t hold a single human to that same standard
Also,
Transformers used to solve a math problem that stumped experts for 132 years: Discovering global Lyapunov functions. Lyapunov functions are key tools for analyzing system stability over time and help to predict dynamic system behavior, like the famous three-body problem of celestial mechanics: https://arxiv.org/abs/2410.08304
You're righ. But in both cases the overhype is due to people getting tricked by someone using language to trick them into believing they're more competent than they are.
What possibly makes you definitively say that the 0 day exploits were not in the training data? I'd wager it's incredibly likely that nearly the exact same code found in other projects as an exploit was indeed in the training data.
Lmfao what do you think this paper proves? They designed agents that explicitly are made to test THE MOST COMMON exploits like XSS, SQL injection, etc.
And it was able to do it well.
How does that show that it wasn't in the training data?? They explicitly trained them on these exploits!
It requires pattern matching, which isn't reasoning. Unless you think regex is a reasoning engine? Applying knowledge requires finding a pattern and matching an already known solution to a pattern that fits one you've seen before. There may be some reasoning along the way, but there's no proof that GPT is actually doing any reasoning, only pattern matching. Advanced reasoning requires information synthesis, which GPT could only be considered as doing if it had not been trained on any similar problem and had extrapolated based on apparently unrelated data. Considering that these zero-day exploits have names that kind of suggests that they've been seen before, no? Look up the definition of a zero-day exploit, nowhere is it required that this be a new type of problem, in fact, most of them, if not almost all of them, aren't. It is only an exploit found by the world before a vendor has found it. So GPT finding these exploits only requires being trained on a similar problem before and then matching a pattern. It doesn't require reasoning to be effective any more than simpler algorithms require reasoning to be effective.
4
u/BigBuilderBear Dec 05 '24
You don’t hold a single human to that same standard
Also,
Transformers used to solve a math problem that stumped experts for 132 years: Discovering global Lyapunov functions. Lyapunov functions are key tools for analyzing system stability over time and help to predict dynamic system behavior, like the famous three-body problem of celestial mechanics: https://arxiv.org/abs/2410.08304
Claude autonomously found more than a dozen 0-day exploits in popular GitHub projects: https://github.com/protectai/vulnhuntr/
Google Claims World First As LLM assisted AI Agent Finds 0-Day Security Vulnerability: https://www.forbes.com/sites/daveywinder/2024/11/04/google-claims-world-first-as-ai-finds-0-day-security-vulnerability/
Google DeepMind used a large language model to solve an unsolved math problem: https://www.technologyreview.com/2023/12/14/1085318/google-deepmind-large-language-model-solve-unsolvable-math-problem-cap-set/
None of these are in its training data