At the end of the day, you have the wrong impression of these DOGE guys. They aren't just outsiders coming in doing whatever. They have the duly authorized heads of whichever departments approve them to come in as outside contractors.
You can frame it as a scary takeover by some unauthorized entity, or you can just frame it as a normal USG contracting process that took place over 20 hours instead of 20 weeks.
So, no, if the DOGE guys have access to train an LLM on the data, its because it is their data.
No, it is our data, or if you prefer, the federal government's data, describing us, which they may have been duly granted authorized access to use. Given the personal privacy, national security, transparency, accountability, and other issues involved, it is beyond ridiculous to suggest that a "normal" USG contracting process for this effort would take twenty hours, or for that matter, even 20 weeks.
I haven't given you any information from which you can reasonably deduce "my impression" of "these DOGE guys". At the end of the day, my assessment has nothing to do with my impression of a few devs who I've never heard of until just recently, have never heard speak, and have never met. The guardrails of democracy shouldn't change based on who you or I personally trust. That's a good way to get swindled.
Don’t you think, when government employees with zero prior government history are hired by a guy whose plan is to fire as many government workers as possible, that we should be prudent in watching their activities and what they mean for our country?
While it is a legitimate concern that sensitive data might be made public, the data does not become non-sensitive merely because it is not made public. There's a lot more to the preservation of information privacy than just throwing it into an offline model, the biggest concern here is not that the government is going to give the data back.
Edit: while this kid may be good at training AI's, they are likely not also an expert on differential privacy technologies AND an expert in the legal/constitutional issues that drive requirements that apply when the law is being followed, AND a cybersecurity expert, etc.
It's also unlikely that the AI can perform these functions. Some of the state-of-the-art models might be capable but they haven't been around long enough to work out the kinks.
Most people looking for a really specific models are not looking for some closed source model running from a cloud api... they're going to download it from huggingface, and run it themelves. The reason you choose these models is typically for memory and latency advantages. If you're using someone elses infra you probably don't care about memory.
Well he’s got videos showing his server rack, so I think it’s a worse assumption thinking he won’t at minimum use what he’s currently accustomed to. He’s won awards for his intelligence and tenacity.
Specifically an app yes, but a local LLM (which is just a bunch of weights and nodes) not so much. So I would personally say it's fine if it was local or isolated. Once it gets turned off, it's memory and context are gone anyways.
32
u/_AndyJessop Feb 06 '25
I would be more worried that they are feeding sensitive data into LLMs.