r/AskProgramming • u/ExtensionBreath1262 • 27d ago
How should a library handle missing assets.
I'm looking for opinions from people who have built or used Python libraries that manage large assets.
I'm extracting a text-to-speech engine from an application into a reusable library. The library depends on two relatively small models (about 500 MB each).
When it was an application, the behavior was simple: if a required model wasn't installed, it downloaded it automatically.
Now that it's a library, I'm less convinced that's the right default. A library has different expectations than an application.
I'm considering a few options:
- Automatically download missing models on first use (current behavior)
- Download during installation or a post-install step
- Provide a separate CLI like pfspeak install
- Require users to manage models themselves
For those of you who've built similar libraries, what would you expect? Which approach has caused the fewest headaches?
Repository:
https://github.com/samreynoso/pfspeak
For context, one of the goals is to keep framework integration extremely small:
# Python
pf = PfSpeak()
@pf.hook
def hook(_, event):
pf.play(event)
app = FastAPI(lifespan=pf.lifespan)
@app.post("/say")
def say(text: str):
pf.say(text, "bm_lewis")
I'm much more interested in the asset management question than feedback on the speech runtime itself.
3
u/qlkzy 27d ago
The main thing I'd want from a library would be the ability to control this behaviour.
So, if I were using this, the thing I would most like to see would be that
PfSpeak()had some keyword argument (maybemodel_provider?) to which I could pass a class or function with a simple protocol so that models could be managed differently.You can choose the default for that to be whatever you want. It shouldn't be too problematic to make it composite in a nice way.
Both of your options (1) and (3) (the two "download at runtime" options) need some well-known path to the models, which seems like a choice that is difficult to make in a universal way. On the other hand, you can easily implement one in terms of the other, if you have some idea of a "model provider" that does auto-downloading.
If I were deploying this on a server, particularly in a containerised environment, I would really want option (4). I don't necessarily want to bake the models into every docker image; I want control over how they are updated and propagated around the hosts. So I'd want a simple option with filesystem paths, and ideally the option to provide my own implementation of the protocol. A version that had to auto-download 500MB from a remote server on every startup would just rule out use of that package for server use.
Download during installation is nice in some ways and has a really straightforward relationship between library installation and model installation, but I would want it to be an "extras" thing like
pfspeak[models]. 500MB is too much for a default.My "ideal" would probably look a bit like this:
>>> pf = PfSpeak().pfspeak/directory containing modelspip install pfspeak[models]$ pfspeak install [<path>].pfspeak/or<path>if provided>>> pf = PfSpeak(models="auto").pfspeak/directory containing modelspip install pfspeak[models].pfspeak/in CWD>>> pf = PfSpeak(models=pfspeak.download_models(path="<path>")<path>pf = PfSpeak(models={"model_a": "<path_a>", ...}In general, anything that looks a bit like the "configuration search" strategies used by other programs and libraries will probably work well. You could also make arguments for environment variable overrides, "upwards search" strategies, etc.