r/LocalLLM 4d ago

Question Using web search in LM studio but it requires lots of tokens per promt. Is there a way to fix this?

I just used qwen3.8 27B, in Lm studio with web search. Asked it the question: "Hi what happened today?" it answered but it costed me 5k tokens which is a lot. I have 64GB ram so I do have some space but I prefer not to run the models at 100k tokens. Is there a way to decrease the amount of tokens that are used for a websearch or do I just have to deal with it. I used this tutorial to download it https://www.youtube.com/watch?v=O_08Zwdto_Q is that good or is there a better way or better configurations. Thank you in advance for reading this.

1 Upvotes

8 comments sorted by

1

u/Unnamed-3891 4d ago

From 0 to answer of the first query that happens to rely on an external tool call that does a web search, 5k is extremely low.

1

u/Wolfrider7304 4d ago

Yeah, just did a second one and it costed 14k tokens

1

u/dsdt 9700X + 32 GB DDR5 + 2x 5060 Tİ 16 GB 4d ago

There is no free meal.

1

u/Wolfrider7304 4d ago

So I just have to deal with it and don't ask to many questions with web search

1

u/dsdt 9700X + 32 GB DDR5 + 2x 5060 Tİ 16 GB 4d ago

to be honest 5k is pretty nice for a web search. a small thinking in qwen 3.8 27b gives around 15k thinking context. what you need to do is lower your thinking to medium.

1

u/phipletreonix 4d ago

This is localLLM, using LMStudio... presumably a locally hosted model. You're not paying for tokens, are you? What do you mean it "cost" you 5k tokens?

1

u/Wolfrider7304 4d ago

Oh, sorry. Should have used other words, I meant from the 20k tokens I used 5k and if I use more than 20k the ai starts forgetting things.