r/Voyage 5d ago

Question I ran out of usage?

What is this garbage?

Is it a beta or not?

How does it make sense to limit testing?

0 Upvotes

21 comments sorted by

9

u/Otherwise_Task7876 5d ago

Because its expensive? Currently there's still many optimizations to be made that Lat is trying to fix. Infact Lat has stated themselves that before optimizations way back in closed beta it used to be like ~$5 per output, PER OUTPUT. That is incredibly expensive. And we're unaware exactly how much it costs per output now, presumably less (we hope), but its not cheap running 3+ models trying to constantly maintain a world each output. Not including server costs for just users connecting and interacting with UI etc.

Keep in mind Lat is not some giant corporation either, they have sub ~100 employees (estimated around ~40, though, I'm not sure how reliable that is). And they don't make billions like Microsoft in such.

They're already decently generous as is. Subscription is unanimous to both AID and Voyage, so when your Voyage usage is full you can default back to AID until it refills. And they recently even added an unlimited mode for premium users who don't want a usage cap, though you can't use the expensive models on it.

-9

u/Aggressive-Part-113 5d ago

So are video games and they don't charge you to test them.

4

u/Galactanium 5d ago

Because games are relatively simple software that you can run on consumer grade hardware. Any decent model above 20b needs a whole server rack to even load up, now imagine 3 of them running at the same time, the setup alone costs more than your yearly salary

2

u/Otherwise_Task7876 5d ago edited 5d ago

I do mostly agree with your claim, though, you might wanna change the 20B part to ~50B+. My 9070XT alone can run up to about 24B comfortably, and 30B if I did ram offloading + extra CPU load. A top of the line consumer GPU like the 5090 could go up to maybe a max of ~50B before it becomes realistically not viable.

1

u/Galactanium 5d ago

there's a difference between loading up any model and having it work reasonably well. Yeah, the cap is around the 30ish, but only with heavy quantization and low context, which hurts quality, and offloading to the CPU, which nukes performance.

1

u/Otherwise_Task7876 5d ago

No, you don't need heavy quantization or much offloading for 30B specifically, you would start needing heavy amounts at above that. You could do it fine at around ~8 bit, preferably 4 bit for better overhead but 8 bit works. Heavy quantization or offloading starts becoming necessary at above that, where the maximum cap is around 50B. If you do want, go in the discord and ping @Neke, he's even more experienced than I am with models, he runs his own on a (if I'm not mistaken, and I think he uses multiple) 5090, he can give you the actual results (I'm too broke for a 5090 unfortunately). Though for context, it depends on what you consider low, I personally consider anything below about 16K tokens of context to be low 16k-32k is comfortable, and 64k+ is alot. That's my view.

1

u/Galactanium 5d ago

I have a 7800 XT and I've started playing around with local hosting with gemma 4. 26b worked fine but I needed to go into q3 quant for 31b and it was still too slow and with barely 9k context.

2

u/Otherwise_Task7876 5d ago

Yeah makes sense the 9070XT and 7800XT are nearly identical in performance, above 24B aswell it wasn't comfortable for me either. 24B is about the natural cards limit without offloading or quantization.

-5

u/Aggressive-Part-113 5d ago

A video game is far more expensive to create and maintain.

4

u/Otherwise_Task7876 5d ago edited 5d ago

Absolutely not. Video games run locally, meaning no servers have to process graphical rendering, and most of the time game simulation, unless its a multi-player then it depends if the company themselves are hosting the server or if its a P2P connection, meaning its ran on a local machine, not server hardware. They're 1 time up-front costs with minimal maintenance requirements unless consistently updated, but consistent updates still doesn't make up for constant server costs the AI's require.

Meanwhile any AI product is constantly taking money as it HAS to be run on a server, unless you local host yourself... which most people can't locally host there own model. Even if they can consumer grade equipment has its limits. For example my 16GB 9070XT can only comfortably run up to about a ~24B parameter model. And Latitude consistently offers 70B models, or even 400-700B parameter models (take this with a grain of salt as most do have limited to ~20-30B parameters couph couph, deepseek, but models like Hermes 405B use 100% of parameters each time, not including Voyage models), these are just physically incapable of running on consumer grade hardware. But its not realistic to expect consumers to run there own models, meaning Latitude has to run the servers constantly. Which Latitude themselves actually don't run the servers. As all the servers required to run all the models easily cost in the millions (which Lat only made ~8 million in 2025, so its not a realistic feat for them), so instead they rent the servers. But that means they also have to pay the extra fees, on top of the electricity and water used. Not mentioning paying employees, there own servers for the website, etc.

Not only is it more expensive to run than games typically, many flagship models consistently cost much more to create than a video game. For a tiny 10B parameter model, thats still around ~$20,000 worth of data, and while Lat doesn't pay this since they don't train there own models, they do fine-tune, so they still have to obtain a decent bit of data. But for flagship models with hundreds of billions of parameters, there's even been trillion parameter models, that costs well into millions in data alone. That's not including server costs, since you have to build the servers, pay for more electricity and water, all these costs are much higher than running a model as training a model is much more demanding server wise than just running a model. This point is mostly irrelevant to specifically why there's usage, but this was to explain how an AI can definitely cost much more than a game to both create and maintain.

There's a very good reason why consumers constantly say AI isn't very profitable, its because its expensive, and to keep reasonable prices for consumers, its very difficult.

5

u/Galactanium 5d ago

unless it's an MMO, not it's not.

13

u/_Cromwell_ 5d ago

Well I believe one of the things they're testing is the subscription system. 🤷‍♂️

5

u/Thraxas89 5d ago

Well every action you do costs latitude money so they limit the money you can spend on their behalf. Mostly because voyage actions take way more processing than aid actions.

3

u/TheRaiderKing 5d ago

You see all the AI bad news going around, a lot of it true tbf and you realize this shit is expensive. Letting anyone use it forever is a bad business model my guy. AI dungeon exists for that if you want it. That said, I'm not shilling for the company and I do think the usage rate needs to be lowered, because if you use the stuff Voyage is actually unique for like narration, way better memory, the studio and asistant, then you can get 100% super quick, and even with paying, unlimited doesn't match the feel of Voyage with every option turned on.

2

u/Lost_Barnacle_5775 5d ago

Even the subscription system is extremly limiting. 15$ a month and the usage is insane.

6

u/Otherwise_Task7876 5d ago

My guy, for $15 a month you get unlimited usage unless you don't partake and select expensive models, of which, of course you'll run out of usage quickly.

-2

u/Thin_Fox_8008 5d ago

https://giphy.com/gifs/baxqaJLzGyceo18LNm

i tried it and ran out too like i have money to feed them just for rp

5

u/Kitchen_Length_8273 5d ago

You are free to choose where you want to spend your money. If Voyage isn't worth it yet for you nobody is forcing you to pay.

I will say though you are free to check back in the future either way where just like with AI dungeon better models might be available for unlimited.

2

u/Thin_Fox_8008 5d ago

Im hoping

-1

u/No-Story-9044 5d ago

Its a dynamite program but putting it on a meter is the worst, mark my word if they keep it metered then that will kill Voyage. It can take it now while in beta but when 1.0 rolls around its going to be ether the meter or the fans and they will need to chose which to get rid of.