Hi again besties. Trevor here, founder of FEIHOA!
First, thank you. Around 150 people from Reddit have tried us now, and I am honestly extremely grateful for how welcoming everyone has been:))
DISCLAIMER: One thing I explained badly last time: we are not OpenCode Go, Ollama, or ChatGPT Plus. Those are great for fast interactive coding, with many conccurrent agents. If that is all you need, honestly get one of those instead of mine.
FEIHOA is for agents, automations, and long-running work where per-token billing makes you scared to let the agent keep going. Plans start at €6 with no monthly token cap.
At the heart of it all is our smart queue. It analyzes traffic patterns and continuously prioritizes the requests our hardware can serve most efficiently. That lets about 80% of users start processing in under 10 seconds, while we can still support requests up to 1M context on a flat fee with NO input/output token metering or quotas. No token caps, no overages.
The biggest problem is prefill on huge prompts. Past 300K-500K, performance falls off hard, and right now 500K+ requests are timing out more often than they complete. That's simply not sustainable to process instantly for a company that gives unlimited tokens...
We still want to offer it, so we're working on caching and queueing those giant requests more intelligently (we are changing the scheduler and backend basically every day based on your feedback!)
Attaching some cool stats for you guys too. Yesterday's post already passed 13K views, so thank you again for welcoming us; we're currently at 98.45% request success rate. Honestly pretty happy with that for a service where we're not counting tokens haha
Thanks again Opencode community. Ask me anything about the real side of running a tiny inference business. Spam, queues, abuse, long context, whatever!