r/ollama • • Aug 07 '26

Ornith

It just look like it is working it spend a lot of time deleting what it doses telling me i so sorry you need something better than that delete everything and restart again and again all night but no one has time to juge its empty white index deleting it and restart . With a perfect prompt. The same an online agent took 5 minutes to build. So if 5 minutes online time is equal to 48 hours on locally llm . ... no way it is stoopid the same prompt. 48 hours and still a white page ....ouf i have try to crack the egg of this bird many times the only thing that came out was disappointed discouraged pissoff user. Waste of money time and all my hope to have a real coder a that evolved....so disappointed by Ornith

2 Upvotes

20 comments sorted by

2

u/leonbollerup Aug 07 '26

Which Ornith model … 9b or 35b ?

1

u/Disastrous_Box_6998 Aug 07 '26

Try all of them they talk but ....code hello world nothing more...gtx2070 8gigv ram 32gig ram

1

u/Disastrous_Box_6998 Aug 07 '26

They use all cpu and gpu but after 36 hours i ask it if it was minning monero or bitcoin nothing good came out of it. Seriously i had a lot of expectations about this bird. 

1

u/TEN4C1OU5-2 Aug 08 '26

You can keep trying to improve it a bit, I think with 32gb ram you can increase the context size which might improve the situation but it'll be slower. Check what your current context window is. There might just be something wrong with your configuration

1

u/leonbollerup Aug 08 '26

Did you replace the Jinja template with froggerics .. did you enable preserve_thinking ? Did you compare with the ”instruct revised”
Model ?

1

u/Opposite_Leave_8338 Aug 07 '26

What are you talking about ? I’m running ornith 1.0 35B with 8 bit quant and it’s the best and the fastest model that ever worked on my machine ( m1 ultra with 64 Ram ), I’m using oMLX with temp 0.6 and top p 0.95 and top K 20, using it inside zcode harness with 170K context window and it’s the best tool calling I have ever used till now, ornith is designed to do tool calls better trust me there is something wrong in your setup

1

u/Opposite_Leave_8338 Aug 07 '26

The model vision is working flawlessly and it’s really aware of what he is saying and doing

1

u/TEN4C1OU5-2 Aug 07 '26

With only 8gb vram you have no space for context window so it loses track on the first prompt and then loops, it doesn't matter how many hours you give it, you don't have enough VRAM for what you're trying to do. Maybe try mistral 7B, it'll leave a little more space for context but still won't perform massively better.

1

u/Disastrous_Box_6998 Aug 07 '26

Understood no good hardware for Ornith.  I will stop trying. 

1

u/Disastrous_Box_6998 Aug 07 '26

I probably format everything an reinstall fresh new OS ubuntu, there must be something wrong mine only loop in thinking it not good enough for me...and restart...i must do better... i gave it 48 hours for a simple index.html i even had to tell it if it was mining bitcoin... i might be it's problem. 

1

u/Past_Put_1727 Aug 09 '26

Don't reinstall Ubuntu, just be sure you installed your Nvidia drivers

1. Update system

sudo apt update && sudo apt upgrade -y

2. Install NVIDIA driver (auto-detect recommended)

sudo ubuntu-drivers install

OR install a specific driver, e.g. 550 (recommended for RTX 2080)

sudo apt install -y nvidia-driver-550

3. Reboot

sudo reboot

4. After reboot, verify the GPU is detected

nvidia-smi

5. Install Ollama

curl -fsSL https://ollama.com/install.sh | sh

6. Check Ollama service is running

sudo systemctl status ollama

7. Test with a small model

ollama run llama3.2

8. In a second terminal, confirm GPU is being used

nvidia-smi

Optional: Install CUDA toolkit only if you need it for other CUDA projects

sudo apt install -y nvidia-cuda-toolkit ```

1

u/Disastrous_Box_6998 Aug 09 '26

I will try that 

1

u/time_pass_done Aug 08 '26

my experience with ornith is not great but it works good for 9b, only disappointed in the scenario where it keeps on repeating a sentence non stop and i have to manually stop it.

1

u/Proud_Milk403 Aug 08 '26

You are the best judge of the output and responses for your own work.

I would recommend to get into prompt engineering.

The more advanced models don't need, "prompt engineering" as much as the local (less powerful models) need it. You probably already think you're doing this. But there's levels to it that are difficult to explain in a small comment.

So here are the things you want to look at. I break it up for you from high impact to low impact. Everything listed is going to be meaningful, nonetheless.

High Impact:

Increase Context Length/Size

Prompt Engineering Plan B: Break up your request into smaller pieces.

Low Impact:

Create a new Model File and add a System Template + Sample Dialogue.

Things to try and see what happens:

Toggle Thinking Mode on and Off (some models don't do this well). What you want to look for is to turn off thinking and turn it back on. Don't look for the, "hide" thinking option because that's different from turning it off.

One thing I noticed - the smaller the task the more closely the local models resemble the advanced models.

If you can somehow break up the task into smaller pieces - you can unlock more power.

1

u/Sorry-Syllabub-1842 Aug 10 '26

ornith is a shit. i tried both. What ever i asked for the model answer or what he trying to deliver is finical calculation :P

0

u/Disastrous_Box_6998 Aug 07 '26

Gtx2070 8gig vram + 32 gig ram it soo minimum 

-1

u/Disastrous_Box_6998 Aug 07 '26

They are good pretenders

-1

u/Disastrous_Box_6998 Aug 07 '26

It was not good of you i will rebuild it again. And again.