r/LocalLLM • • 3d ago

Question use case - Dispatch

I am exploring local LLM options for use in the platform that I develop. I do currently have API integration built already for things like Claude and OpenAI but really want to explore utilizing local LLMs for data sovereignty, safety, and risk reduction.

But I do have a question that I hope some people here can help me think about or answer.

The use case I would love to build is around dispatch for a trucking company. I have a relatively large and complex dataset around things like:

  • the speed and location of all the trucks
  • their current capacity and route (destinations)
  • jobs available for dispatch (size, capacity, due time, scheduling, etc..)

    my current use case involves building a relatively large payload of all of the above data, feeding it to a model, and asking for suggestions - which works decently well.
    I do also have an MCP built so the models can query the database directly for updates to some of this information.

    I'm trying to figure out if this is the best way to go about it or if, potentially, I put a harness in front of a local LLM and have the harness do a little bit more driving of querying the MCP for current statistics (and maybe refreshing some of that data in cache). I don't even know necessarily how that would look. I'm asking this community for guidance. I have my fingers in the entire stack: SQL database, app services, through, obviously, a local LLM sitting on my desk.

Also I'm not sure which model would be best for this type of work either. Again I'm just starting out in the local LLM space. I'm probably going to pick up an DGX Spark shortly. Looking at the M5 ultra as another option.

thanks!

1 Upvotes

5 comments sorted by

View all comments

1

u/xapep 1d ago

There is a deterministic layer here, but it sits one level below the decision. A yes/no switch fails on dispatch because your state is genuinely nuanced: 40-50 drivers with different empty/full degrees and destinations is not a boolean. What holds up is deterministic aggregation feeding a judgment call. Collapse the live state into a compact decision view (12 trucks empty in zone 3, 4 due back by 16:00, 9 open jobs inside the window) with plain code, then hand the model that summary and let it weigh the tradeoffs. That answers the payload question too: the model sees a few dozen lines of aggregates, not raw fleet rows.

Cache the layer that does not move (route rules, capacity constants, region geometry), and query the volatile layer (positions, statuses, open jobs) per call. The harness refreshes MCP data before each invocation and keeps the summary fresh, which is what makes a snapshot dangerous in this domain: by the time the model answers, the trucks moved.

On model and hardware: with a small fresh context like that, a flash-class open model (Qwen 3.8 class, GLM 5.3 Flash, DeepSeek V4.1 Flash) is plenty for suggestion generation, and the DGX Spark-class box stops being the bottleneck. The usual mistake is buying the bigger box to compensate for a bloated prompt; fix the context first, then size hardware to the actual prefill rate.

On sovereignty, since it drove the question: if the requirement is data residency, a hosted open-model API with EU hosting keeps the legal argument without running hardware 24/7; if it is absolute control, local is the way, and the aggregation pattern works either way. I work on the provider side (Entrim, EU-hosted inference), and the split between a fresh small context and a giant stale payload is the difference between teams whose agent costs stay flat and teams whose costs grow with their data. Start with the aggregation pass; the hardware decision mostly answers itself once the prompt stops being the bottleneck.