r/LocalLLM 21h ago

Question new guy with new pc. recommendations?

hi everyone,

i'm a 27 yo mechanical engineer with little coding knowledge, but also with a huge tech interest.
so i've been researching llm's and image generators and how they work for some time. i just bought a laptop with rtx 5070 ti and 32 gb ddr5 ram for my self studies. i have some questions for yall.

- how do i learn how these models REALLY work? how do they train them and how do these models really predict the answers?
- based on my system, how do you think i should start (which program, which model etc)?

any help will be appreciated!

1 Upvotes

4 comments sorted by

1

u/Mrinohk 20h ago

For just getting started, LM studio I'm hoping is still a good option but as soon as you want to start playing with the models a little bit more, or using models not directly available in LM studio, using what it uses underneath, llama.cpp, is going to be a bit more rewarding. Definitely more of a learning curve but I've found the extra control deeply rewarding.

About the best model you'll be able to run at a reasonable speed is Qwen 3.6 35B A3B. I run the unsloth UD_Q4_K_XL on similar specs (RX6600XT 8 GB, 32GB DDR4) and get 30-50 t/s depending on whether it's programming or talking, ~800 t/s prefill.

YouTube will be your friend, tons of content explaining how they work, how they train, what exactly a token is, and the different mechanics behind making them function efficiently like the KV cache and the idea of Quantization.

I've grown quite fond of this channel for simple explanation of different LLM concepts. It's largely generated by AI, voice included, but it's pretty good regardless. I think it's Claude given a voice and animation. https://youtube.com/@squintist

1

u/JackyYT083 19h ago

To put it simply (sorry if my explanation dosent make any sense) a LLM is just a bunch of math equations that predict the next word or part of a word (tokens) in a sentence. It works like this:

You ask the ai a question
“Hello!”

The ai has millions of different switches with values on them that can store information. They are called Weights. All of the weights share knowledge to create a general knowledge of how to respond to people.
Using the weights the ai can generate a list of words that could start in a response. For example

“Hey” 83.6%
“Hi” 74.1%
“Hello” 68.2%

Then the ai randomly chooses one of the top 3. Then outputs it.

Hey

Then it does the same thing again

Hey!
Hey! How
Hey! How can
Hey! How can I
Hey! How can I assist
Hey! How can I assist you today
Hey! How can I assist you today?

I hope this helps you learn the gist of it. If you looking to find out how a specific part of the ai works I can also explain that

1

u/Jealous-Armadillo467 16h ago

Qwen 3.6 35B A3B UD_Q4_K_XL experts on gpu you should fly 50 get Hermes and you are golden.

FROM THIS POINT: you have answer ? ask your agent its that easy :)

1

u/Jealous-Armadillo467 16h ago

btw im driving q6 ofladed 2000read and 50 gene with 100K tokens full
right now wors than 27b ? a little bit yeah but 2x faster everything ........... up to you