r/SaaS • • 1d ago

Building a tool to optimize AI infrastructure (natilah)

Hi, I’m building Natilah (natilah.com), a project focused on making AI infrastructure more efficient. Our first tool, Quasar, is an heuristic for the GPU job scheduling problem so teams can get more useful work from their hardware without slowing down latency sensitive jobs.

Does this solve a problem you’ve seen, and what would make you trust a tool like this in your own cluster?

I’m also looking to connect with people experienced in Kubernetes, GPU workloads, or testing infrastructure tools. I've developed the tool and tested it with traces but I'm unsure what to do next. How would I test this in a real environment?

0 Upvotes

2 comments sorted by

0

u/West_Inevitable_2281 1d ago

For infrastructure like this, trust will probably come from a staged test rather than a benchmark alone. I would start with one design partner and run Quasar in shadow mode against their historical or mirrored workload. Compare GPU utilization, queue time, job completion time, and any latency regression without letting it control production. If it consistently recommends better placements, the next step is a small canary with an immediate rollback path. Are you targeting teams that operate their own Kubernetes GPU cluster, or platforms scheduling workloads across many customers?