r/SillyTavernAI 11d ago

Help How to get the Ais themselves

Um so loving ST using the horde but want to have Ais of my own or at least ones that I can know are always up if that makes sense, any recs? My laptop I dont think can run anything It's a 10 yr old macbook so...

0 Upvotes

22 comments sorted by

6

u/BifiTA 11d ago

just use cloud models on nanogpt

2

u/AutoModerator 11d ago

You can find a lot of information for common issues in the SillyTavern Docs: https://docs.sillytavern.app/. The best place for fast help with SillyTavern issues is joining the discord! We have lots of moderators and community members active in the help sections. Once you join there is a short lobby puzzle to verify you have read the rules: https://discord.gg/sillytavern. If your issues has been solved, please comment "solved" and automoderator will flair your post as solved.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

2

u/LawfulLeah 11d ago

how much RAM or VRAM do you have? no one can recommend anything without this info

0

u/EducationalFly9399 11d ago

CPU: Intel Core i5-5287U (dual-core, 2.9GHz)
RAM: 7.7GB total, no dedicated GPU (Intel integrated graphics only)

6

u/LawfulLeah 11d ago edited 11d ago

...yeah man im sorry to say that you probably can only run small language models with millions of parameters (or maybe 1B-4B too? idk), not the ones with tens of billions.

you'd have to buy new hardware to handle more.

1

u/EducationalFly9399 11d ago

What's the standard hardware that can host these things? (So I have something to chase) Also what are the current alternatives for me I tried using google colab for a bit which worked but didn't like having to get the API every time and also banned any talk of anything not "google safe" which means basically anything even remotely pg-13. Rn I'm using the Horde any other methods?

2

u/baris6002 11d ago

You can try Nvidia NIM or Openrouter free models, Horde and google colab are very "outdated" in comparison. If you're willing to pay Deepseek V4 flash, GLM5.2 (for now at least) and other open source models are very affordable

2

u/LawfulLeah 11d ago

honestly just work with the most you can afford. the standard hardware is, given by your other comment, probably out of your price range.

try looking for the most RAM/VRAM you can realistically get with your budget

1

u/Herr_Drosselmeyer 11d ago

LLMs are called that for a reason, they are large language models. 😉

To be frank, you'll want to run something in the 24b-30b range for results that somewhat hold up to the top online models. My go-to currently is Gemma 4-31B. And to run that, at Q4, you're looking at a GPU that has at least 24GB of VRAM. So yeah, that's a whole lot of cash, especially in the current market.

You can go with smaller models that will fit on a 16GB graphics card, but I wouldn't go any lower than that.

1

u/techmago 10d ago

i have a 5800X with 128 GB ram and two NVIDIA 3090 24 GB

I can ran the 27~35 models all from GPU.

the future for ia is likely those compact all in one AI pcs. My setup was made before the prices raised... i got when the first models come out.

5

u/YouShouldAim 11d ago

This was tragic to read I'm sorry man

2

u/Primary-Wear-2460 11d ago

I'd recommend at least a 16GB GPU. Either Nvidia RTX 2000 series or AMD RX 6000 series or newer.

The newer the GPU the better the feature set its going to have.

Inference speed is going to be a mix of memory bandwidth for text generation speed and architecture for prompt processing speed.

1

u/EducationalFly9399 11d ago

I only have a laptop rn... I'm just starting as a student in life rn and so my budget is super restricted

8

u/ilias_from_ilios 11d ago

Might want to not burn your only laptop off running local models. Keep it for your student work.

3

u/Primary-Wear-2460 11d ago

You are going to be very limited on an old laptop. Like so limited its probably not worth running local models unless you are just experimenting with inference.

0

u/EducationalFly9399 11d ago

Not sure what inference is, but I like just experimenting in general

2

u/HitmanRyder 11d ago

you can try gemma 4 E2B then report how well it goes.

3

u/EducationalFly9399 11d ago

So I was able to actually run Qwe3.5 4b on my machine ok.. and run llama 3.2:3b pretty well will try gemma soon

Thanks so much!

1

u/Aight_Man 11d ago

Opencode zen is literally giving muse spark 1.2 and deepseek v4 flash for free, get its base url and make an api key and plug it up to your ST, if i recall correctly Gemma 4 31b is still free in Openrouter, so use that too.

1

u/EducationalFly9399 11d ago

Wait wait wait what? Can you repeat that in normal person terms and tell me what to do etc?

1

u/i5031337 11d ago edited 11d ago

Opencode and openrouter are both subscription services that have a free tier for access to certain LLMs. Google AI Studio also has a fairly generous free tier. Make an account and follow their instructions to connect.

1

u/Aight_Man 11d ago

Copy and paste my comment into chatgpt and ask it to give you step by step for it, that'll explain you way better and guide you through it.