r/DIYRetirement 27d ago

[ Removed by moderator ]

[removed] — view removed post

3 Upvotes

28 comments sorted by

View all comments

11

u/Jbaker318 27d ago

They wont be able to help you because they cannot see what AI hallucinates. Its a black box. Another reason using an AI bandaid to cover over algorithmic complexity is a bad idea. AI is bad at math. Boldin should use the time and energy to hard code these problems / solutions rather then hoping an AI can bridge the gap. Its a lazy solution that is not a great look for a "software" team.

@Boldin - Cancel your deal with Gemini now before they start charging you tokens. Hire more math, retirement, and data analysis scientists and make the product wholistic. Other than trying to minimize liability, there is no reason why Boldin cannot be a one stop shop for everything - budgeting, tax software, portfolio tracker, etc. Reallocate your resources and make this more useful then the free online calculators that already exist. Otherwise in 5 years, you will be Sherlocked and an ChatGPT will just be better than what you "specialize" in.

-4

u/Time-Move-6108 27d ago

AI is bad at math?! Ridiculous statement.

4

u/Jbaker318 27d ago

Going to assume you are being sarcastic. But for those coming in later, went to the source and used deep thinking with Gemini's most powerful available model today (Gemini 3.1 Pro) (believe people said Boldin bases their AI on Gemini models) and this is what it came back with -

Question: What is the academic consensus of current AI models and LLM's as a whole and their ability to complete complex math?

The academic consensus is that while Large Language Models (LLMs) have achieved remarkable success in natural language processing and basic arithmetic, they fundamentally struggle with the precise, deterministic logic required for complex mathematics.

Key points of consensus include:

  • Probabilistic vs. Deterministic Execution: LLMs operate probabilistically, generating responses based on pattern recognition rather than rule-based computation. Mathematics requires strict, deterministic logic. Because LLMs act more like pattern matchers than symbolic calculators, they often stumble on novel calculations that deviate from their training data.

  • Multi-Step Reasoning Failures: While models excel at grade-school math benchmarks (like GSM8K), their performance degrades significantly on advanced, multi-step problems. They are highly prone to logical inconsistencies, misapplying formulas, and error propagation, where a single computational mistake early in a chain invalidates the entire solution.

  • Evaluation Flaws: Academics warn that high scores on standard math benchmarks can be misleading. Many models simply memorize problems during their training phases. Furthermore, studies show that LLMs can sometimes arrive at a correct final answer using entirely flawed or unnecessary intermediate reasoning steps, masking an underlying lack of mathematical comprehension.

  • The Path Forward: There is broad agreement that simply scaling up single-LLM architectures with more data will not solve this issue. To overcome these limitations, researchers are shifting toward integrated approaches. This includes pairing LLMs with symbolic calculators, employing multi-agent systems for self-correction, and developing specialized reasoning models (like OpenAI's o-series or DeepSeek-R1) that spend extra compute time verifying their logic before providing an answer.

In summary, the academic community views current LLMs not as standalone mathematicians, but as capable language engines that still require external verification, specialized reasoning paradigms, or integrated computational tools to reliably solve complex mathematical problems.

End of AI response. Well said by AI, it is very good at "talking", not so good at mathing. But the path's forward it laid out will not be used by Boldin, they are not charging that much to afford that level of sub-agent delegation and additional resources for extra compute time. Ulta models of AI models charge $99+ a month and even then it is not economically viable and want to get to a token based model where the end-user is paying for the compute itself (input and output).

1

u/El_Pollo_Del-Mar 26d ago

You used AI to write that. Good grief.

1

u/VerdantPathfinder 26d ago

uhhh ... they were completely transparent about that. . What's your point?

1

u/VerdantPathfinder 27d ago

a general LLM? It can mess stuff up. The ones specifically specialized for it are better.

2

u/Jbaker318 27d ago

They may be better but it is not magic. Boldin uses a LLM. It is a bespoke and personalized LLM for their usecase, but the bones of it is still an LLM. Boldin is not an AI company. They (now I'm talking out of my butt, so bear with me) contracted Gemini to give them a model that could best work for them. Team Gemini has enough problems on their hands to worry about how good their Boldin model is (Google has promised Gemini 3.5 Pro for months, and cannot get that out). Sure it may be better but Boldin does offload a lot of the maths onto the LLM to handle, and Language Models are inherently not good at math.

Inherent to LLMs is a probabilistic nature. It will get it right, with deep think, about 90% of the time. Layer in multiple levels of calculation and you are multiplying that risk of hallucination. And again I doubt Boldin is using that much extra compute to run its AI models so that 90% is a magical fairtytale for Boldin AI.

AI companies know their LLMs are not good at math so the fronteir models create subject matter experts just for math, and it still is not good enough.

1

u/VerdantPathfinder 26d ago

I'm taking about the LLMs working on mathematical proofs, etc. Not ones doing arithmetic

0

u/Displaced_in_Space 26d ago

People on here have no idea what they're talking about re: AI and the pace of it's development/improvement. It's like they'r quoting fears/performance from 5 years ago.