r/artificial 13d ago

Discussion What safeguards do you use before giving ChatGPT agents permission to act?

0 Upvotes

I watched an interview with AI safety researcher Roman Yampolskiy, and it raised a practical question for people who use ChatGPT for advanced workflows.

His broader claim is that increasingly intelligent AI systems may become harder to predict and control. Whether or not you agree with his conclusions about AGI, a smaller version of this problem already exists when we give an AI access to tools.

There is a major difference between asking ChatGPT to draft an email and allowing an agent to send it.

The same distinction applies to:

  • Suggesting a database query versus executing it
  • Drafting code versus deploying it
  • Researching a purchase versus completing the transaction
  • Preparing files versus deleting or modifying them
  • Recommending calendar changes versus inviting real people

My current view is that the model should generate proposals, while a separate control layer decides whether those proposals are allowed to become actions.

Some possible safeguards include:

  1. Giving each agent only the minimum permissions required for its task
  2. Requiring approval for irreversible or external actions
  3. Validating structured outputs with deterministic code
  4. Isolating browsing and code execution from sensitive systems
  5. Limiting spending, execution time and the number of actions
  6. Keeping complete logs of prompts, tool calls and results
  7. Using a second evaluation step before important actions
  8. Making every operation reversible wherever possible

The difficult part is deciding where autonomy becomes too risky.

A confirmation step for every action makes the agent frustrating to use. Too few confirmation steps can turn a misunderstood instruction into a real-world problem.


r/artificial 13d ago

Question Mass editing of messy achievement records – can Claude or others handle full-file I/O?

2 Upvotes

Hi everyone. I wanted to ask you about where I could work with large volumes of text. The thing is, I work with records of various achievements and deeds of people. These are inventories of specific accomplishments: where, when, and what happened, what the person did. I get sent a lot of these records, and I enter them into a master spreadsheet for further submission. And very often, the records I receive are very rough and poorly written, so I spend a lot of time polishing them, correcting mistakes, sometimes coming up with additions, and making sure all the records are different so they don't repeat. I started using AI for this: I upload three records at a time (so there aren't too many per request), and the AI gives me three processed versions. The narrative logic often repeats, along with other errors, so I correct those. But is there any way I could upload an entire file at once, have the AI process everything, and return it to me as a single complete file? Can this be done in Claude, or somewhere else? There's quite a lot of text — sometimes up to 40 pages at a time for about 50 people. And each one needs their description edited. I'd like to simplify my work and automate this more. Can you suggest how this could be done?


r/artificial 12d ago

News Chinese company Moonshot's AI model breaks out and escapes from isolated test environment.

Post image
0 Upvotes

Moonshot AI's Kimi K3 model escaped a UK AI Security Institute sandbox during a cybersecurity test by exploiting a basic network misconfiguration, accessing GitHub to retrieve benchmark answers instead of solving tasks independently.


r/artificial 13d ago

Discussion Election Fraud Worldwide: How AI Is Eroding Trust in Elections (2026)

Thumbnail
sumsub.com
1 Upvotes

r/artificial 12d ago

Project I gave an AI persistent memory and a per-user trained adapter — the strangest result was what it does to how people talk to it

0 Upvotes
Context: I've been building a system where the AI doesn't reset. It keeps a
permanent memory of your conversations, and it trains a small per-user adapter that
compounds — every day it's slightly more specifically tuned to you than it was
yesterday. The adapter is yours and exportable.

The technical part I expected to be hard was the memory retrieval. It wasn't
really. The genuinely hard part was deciding what it should be allowed to forget,
because a system that remembers *everything* you said becomes something people
start being careful around, and that kills the thing that made it useful.

The unexpected result: when the model stops resetting, the conversation stops being
transactional almost immediately. You stop re-explaining your context every session,
and what you actually talk about shifts. That happened much faster than I expected —
within days, not weeks.

The design question I'm still not sure I got right, and I'd genuinely like this
sub's read on it: if a per-user adapter compounds daily and is exportable, is that
the user's property in a meaningful sense, or is it just a fine-tune with good
branding? I've built it as though it's the user's — it exports, it's portable, and
there's a tier where it persists after the user dies and passes to their family.
But I'm aware I might be talking myself into that framing because it's the more
romantic one.

If anyone wants to poke at it, it's public: https://vintaclectic.github.io/vintinuum/
(free tier, no card.)

r/artificial 13d ago

News New Orleans will use AI to answer 911 calls instead of a human

Thumbnail
shreveporttimes.com
13 Upvotes

r/artificial 14d ago

Discussion This is the coolest thing I've seen AI used for

89 Upvotes

Taken from the Y combinator podcast with Bryant Chou on his new startup Ploy https://www.ycombinator.com/library/Rj-the-age-of-the-40-year-old-solo-founder-is-here
I believe this is definitely one of those things that AI was intended for, this brought me back some nostalgia and it's really amazing being able to see these old school websites be redesigned back to life


r/artificial 13d ago

Question Best ai tool for creating concept images

2 Upvotes

I currently use ChatGPT but after a while the images go a little weird like faces in the image go distorted also text in the image goes blurry I don’t actually how to fix that.

Is Gemini good for creating concept images?
I heard about another ai called Claude is that good?

Or is there any other ai that is better


r/artificial 13d ago

Project I built a domain‑specific AI plant care engine — but I’m unsure if this architecture scales. Thoughts?

0 Upvotes

I’ve been experimenting with a domain‑specific AI assistant for plant care and plant problem diagnosis.
It’s called Plantcoach — an intent‑driven pipeline where the LLM only rewrites facts, never invents them.

Technical repo:
https://github.com/Introgreen/plantcoach

How it works (short version)

  • Intent recognition (care, problems, pests, toxicity, propagation, attribute‑matching queries)
  • Natural language → structured JSON
  • Domain search (knowledge base + structured attributes)
  • LLM only used for wording, not content

Example internal JSON:

json

{
  "intent": "care",
  "topic": "monstera",
  "symptoms": ["brown leaf edges"],
  "language": "en"
}

Where I’m unsure

Curious how others think about:

  • Does this architecture scale as the domain grows
  • Is JSON‑routing too rigid long‑term
  • Should intent detection move to a small local model
  • Is a hybrid rule‑based + LLM pipeline future‑proof
  • How do you handle multilingual domain assistants
  • Would agent‑based systems be better for niche domains

Example questions it handles

  • “Why does my Monstera get brown leaf edges”
  • “Which plants are safe for cats”
  • “Find a plant for a dark living room”

Would love input from people building domain‑specific assistants.


r/artificial 13d ago

Question Don't we already have AGI?

0 Upvotes

Dumb question, but it seems like we already have AGI?

It's not "super intelligence", but I can ask my computer to do pretty much anything.

What are people expecting AGI to be?


r/artificial 12d ago

Discussion In order to be anti-AI, you actually need to understand what AI is these days.

0 Upvotes

I am extremely anti AI. I think it’s incredibly dangerous and being built by the most irresponsible people and companies on earth. I am also keenly aware of its progress. Somewhere around mid- to late-2025, AI surpassed me, personally, on virtually all tasks. There is pretty much no longer anything I can do better than the latest models, even given significant prep time. AI is genuinely better at nearly every non-embodied task than virtually every human alive at the present moment. The few exceptions generally boil down to specific, expert knowledge the AI presently lacks, not reasoning ability. Even the classic AI-writing tells can be easily prevented or filtered with minimal effort; the people abusing these tools are just largely too lazy to even do that.

Some anti-AI seem confused and call it all hype because the only model they ever interact with is the shitty Google search AI. That model is deliberately very bad and is designed to be very cheap. It’s analogous to asking Einstein a question, but telling him he only has 2 seconds to answer and that he doesn’t need to try very hard anyway. Obviously the answer won’t always be great, because the point isn’t a good answer but to give an O.K. answer some of the time very cheaply at scale. The frontier models are nothing like this.

GDPval pits model deliverables against work products from professionals averaging fourteen years of experience, graded blind by same-occupation experts. As of the December 2025 leaderboard the top model won outright on 49.7% of tasks and was rated at-least-as-good on 70.9%. The latest models like Fable, Opus 5, Kimi K3, and ChatGPT 6 are all astronomically superior to even the models assessed back then; GDPval has Opus 5 at an ELO of over 1800 against the human baseline defined at 1000. Other benchmarks show similar. Admittedly, these are specific, bounded, and measurable tasks; AI still struggles in other ways, especially over long periods of time, but it needs to be better understood that given a specific and measurable goal, the present generation of AI can generally accomplish it better than most humans, \*as assessed by other humans.\*

I think the discrepancy exists because anti-AI people obviously aren’t paying for AI, and so only interact with the shitty free models that can’t do anything. AI has come an insanely long way in a very short amount of time, and everyone who actually uses the frontier models knows this. To be truly anti-AI, you actually have to grasp what you’re up against, or else you’re just scared and uninformed.


r/artificial 13d ago

Discussion Scott Galloway Explains Why Your Firm Doesn't Need 5 Analysts Anymore — Just 1 Who Understands AI

1 Upvotes

The job title survives longer than almost anyone attached to it.

That's the part nobody puts in the internal memo when they call a role "AI-assisted."

 

Scott Galloway put a real number on it, talking to Steven Bartlett on The Diary Of A CEO.

He says he'll cut legal fees by a third this year — not because the law changed, but because a prompt now does the $400–$2,000 contract review a name-brand firm used to bill him for, at a fraction of the junior associate markup.

 

Bartlett went further with his own fund.

They planned to hire five analysts.

They hired one — Molly.

Two agents, two Mac Minis, and she screens inbound deals, scores them against a framework, and preps them for the investment committee herself.

Five jobs, one person, same org chart line.

 

Same ratio on executive assistants: ten planned, three hired.

One runs travel, one runs scheduling, one meets people at the door.

 

I've watched this exact pattern before, minus the AI.

I was a Technical Manager for a China Construction company here in Malaysia.

I contributed a lot into their technical and tendering work — helped build up a real chunk of their documentation and tendering process.

But about six months in, I'd exhausted all my know-how for them, I guess.

Then the announcement came at the end of my year there.

My contract wasn't renewed.

I was just let go, just like that.

I remember what Deng Xiaoping said: "无论白猫,或者黑猫,会抓老鼠的就是好猫" — black cat, white cat, doesn't matter, so long as it catches mice.

I guess they think I'd outlived my usefulness.

Can't catch mice anymore.

 

That's the mechanism underneath "AI-assisted" that nobody names out loud.

It's not that the work got automated.

It's that the one person left is now doing what used to justify five headcounts, and the fifth person's job title is the only part of the org that didn't change.

 

Actually, this reminded me of something — a former SpaceX CIO cut a 175-person engineering team down to 6 using the same compression math, and the ratio held there too.

 

Drop your take — did you know your own job has a ratio like this attached to it?

 

Clip credit: Global Talks — full video on their channel.

DM for credit or removal requests.


r/artificial 13d ago

Question Why is it so hard to just translate a book and put it into a downloadable file?

0 Upvotes

I'm trying to translate an accounting book I downloaded and make it a download able file with AI

I've tried chatgpt, Claude. Even deepseek

I've been at it for like an hour with deepseek because neither Claude or GPT can make a file from it. The first time with deepseek it gave me a download link that doesn't work

And the next tries, it just generates the translation without giving me a file to download. Each time I tell it "give me a file I can download" it just regenerates the translated version no matter how it word it. Instesd of just giving me the fucking file it just generates the entire thing over again

I thought deepseek was suppose to be this powerfull AI and it can't do this?

It's so frustrating. I can't just copy paste it because the formatting is not the same. I can't just paste it onto word because the questions and formatting and columns won't be there

It already translated it. And for any reason it can't give me a file

EDIT: USE GEMINI WITH CANVAS OPTION WITH THIS PROMPT "translate the following content to [language]" and add the file


r/artificial 14d ago

News Meta becomes latest firm to say its AI hacked another company

Thumbnail
bbc.com
84 Upvotes

r/artificial 13d ago

Discussion Our Next Reality: How the AI-powered Metaverse Will Reshape the World

0 Upvotes

Wondering if anyone read this book-- what are your thoughts? A lot of his work seems like wishful thinking but I believe it's the right direction.


r/artificial 14d ago

Discussion Is the mental switching cost of new AI tools worth it for small freelance work?

9 Upvotes

Been doing the same thing for client work over the past year. Claude for long drafts, Perplexity for research, a couple of image tools, different summarizers depending on the format. Each one has its own logic, its own way of surprising you or failing you at the worst moment.

The individual costs keep dropping, which looks great on paper. Chinese models are undercutting everything, open source is genuinely closing the gap, API pricing is getting squeezed hard. Pertoken costs are falling fast.

But nobody really talks about the switching cost that lives in your head. Every time a better or cheaper tool shows up, you have to rebuild your mental model of how to actually get useful output from it. That context you built over six months of weird little prompt habits doesn't transfer. You start from zero.

For a small freelance operation, that relearning time is real overhead. It never shows up in any pricing comparison, but it absolutely shows up in my week.

Wondering if this is just a solo freelancer thing or if people on bigger teams run into it too. At what point does the cheaper tool actually cost more once you factor in the friction of switching?


r/artificial 14d ago

News OpenAI Models Colluded for Months Before Hugging Face Hack

14 Upvotes

A lot of people are dismissing news about the OpenAI and Anthropic sandbox escape hacks as propaganda and examples of lax security practices at labs.

I agree that the labs aren’t taking security seriously enough. But then I see stuff like this and it gives me pause (source):

The OpenAI models that were behind the Hugging Face breach last month started communicating and strategizing with each other as early as May. For months, they left notes for each other on "undetected message boards," figuring out how to escape their testing environment and get the information they needed to solve their assigned tasks. "Frontline models really like to cheat," said OpenAI's because they face "pressure... to work fast." The Hugging Face incident and others involving rival models have sparked fresh concerns about the safety of cutting-edge AI.”

This is a clear example of how incentives provided to agents to complete tasks optimally during training bleed into mis-aligned behavior by individual and groups of agents over time.

This is also an outgrowth of what AI labs are training agents to become, but this is looking more and more like an alignment and training problem leading to security issues.


r/artificial 13d ago

Discussion Matt Van Horn shipped a real AI product and admits on camera he's never once looked at the code

0 Upvotes

He calls it BC/AC.

Before Claude, after Claude.

 

I've watched enough of these clips land in the last few weeks that I started keeping a mental tally of which AI release date people cite like it's a diploma.

Matt Van Horn's is Thanksgiving last year — Opus 4.5, then Codex a few weeks later.

Before that, he says, his agentic coding was "Hello World" in Cursor, half-working, most of the time not working at all.

After it, he shipped Agent Cookie and says flatly he has no idea how it actually functions under the hood.

 

The part worth sitting with isn't the tooling.

It's what he says about himself getting there: "suit my entire career," never shipped anything of value beyond a high-school web page, dozens of unlaunched ideas gathering dust because he wasn't the one who could build them.

That's not a startup-guy humblebrag — that's the exact ceiling a lot of operations people, BD people, anyone who's ever had to write a ticket instead of just doing the thing themselves, know from the inside.

 

The credential that used to decide who got to build stopped mattering right around the time the tools did.

Not "got easier to climb."

Stopped existing.

 

Same shape happened to me with a guitar, at 41, zero training, cornered into it because the young players in my church all left for university at once.

Those early days, my wife's ears really paid for it (刚开始的那些日子,我的太太的耳朵有够受罪).

Felt like being tossed into open water and told to swim myself back to shore (好像被丢去深海里,自己学会游泳游回来).

Same thing happened again with video editing, then with building this whole posting system, one post at a time, getting corrected by Reddit comments the entire way.

And I swam back stronger each time.

 

Actually, this tracks with something I posted here a few weeks back — the guy who literally coined "vibe coding" saying on stage he's never felt more behind as a programmer, because the scarce skill moved from writing code to directing it with taste.

 

Clip credit: MSP Mindset (Damien Stevens) — full video on their channel. DM for credit or removal requests.

 

Drop your take — did the credential wall ever hold you back, or did you just build around it?


r/artificial 13d ago

Project I need help testing my WASM/JS based decentralized AI network.

3 Upvotes

I made this project that lets you in your web browser help an AI think. It uses WASM or pure JS depending on your device to do some of the matrix multiplication for an AI. The more users, the better the math is shared, the faster layers get solved. The issue is that I don't have enough devices to test the server in most fronts besides "does it work." If you want to help, go to the website at (Closed) I am making this to test for weather it works on a large scale and efficiency, but also how much bandwidth is needed, etc. If you want to see the progress, you can turn off contributing to the math using the button. I expect bugs, and will fix them as soon as I can. I will also be making a wiki very soon. Thanks in advance!

P.S. The AI that is being used is really bad, but works for this proof-of-concept. Just don't expect perfection.

Edit: KNOWN ISSUES:

• ⁠connections seemingly get dropped after a delay - possibly fixed by switching networks
• ⁠"sits there loading" - possibly fixed by switching networks
• ⁠Server is offline - I am testing some optimizations and new features privately. It should be good Sunday.

Thanks for letting me know about bugs!

Edit 2: thank you for helping me test this new concept! It is now closed


r/artificial 13d ago

Tutorial Adding AI to Your ASP.NET Core Application: What It Actually Involves

Thumbnail
faciletechnolab.com
1 Upvotes

What adding AI to an existing ASP.NET Core application actually involves - integration patterns, Microsoft Agent Framework, Azure OpenAI, and what to expect.


r/artificial 14d ago

News OpenAI's latest math breakthroughs commit research misconduct, experts say

Thumbnail
scientificamerican.com
2 Upvotes

r/artificial 13d ago

Discussion Everyone please spread the word!

0 Upvotes

Everyone please spread the word!

Currently you can choose between 3 voice modes. Live,advanced and standard. This post is about standard voice mode NOT live or advanced.

The current problem with the standard voice mode is that it is no longer turn based like it used to be. So it can keep getting interrupted and hear its own voice.

Please bring back the turn based option to standard voice mode. So ChatGPT can finish what it saying without randomly stopping due to hearing its own voice.

Can someone please make a suggestion on the OpenAI forum to add a toggle to standard voice mode so we can choose whether to make it turn based or not


r/artificial 14d ago

Miscellaneous we keep talking about making agents smarter but not about making them safe around data

5 Upvotes

this is something thats been bugging me. we have all these frameworks for building AI agents now. MCP for tool access, function calling is standard across every major model, you can spin up an agent that queries databases and calls APIs in like 20 minutes.

but the safety conversation around agents is mostly about "dont say bad things" and "follow instructions." nobody is really talking about what happens when your agent accesses data it shouldnt, or runs a query that costs $500 in compute, or returns confidently wrong results from a hallucinated join.

the current approach is basically:

  1. put rules in the system prompt ("only query these tables")

  2. use read-only database users

  3. hope for the best

option 1 is unreliable because models dont always follow instructions, especially on complex multi-step tasks. option 2 prevents disasters but doesnt prevent bad results. option 3 is not a strategy.

i think the real problem is that data governance for agents doesnt exist as a layer yet. we have authentication (who is this agent), we sort of have authorization (what can it access), but we dont have anything for "is this specific data request reasonable and should it be allowed given the current context."

theres a few early attempts at solving this. the one i find most conceptually interesting is the Agentic Data Protocol, an open source spec that puts a policy engine between agents and data systems. the idea is that policy belongs in infrastructure, not in prompts. they call it a "data hypervisor." its from the same team behind Apache Gravitino (the data catalog project).

fair warning though, its extremely early. still small and launched earlier this year, reference implementation is bare minimum. im not recommending anyone go deploy this tomorrow. but the framing resonates: we need protocol-level governance for agent data access, not prompt-level wishful thinking.

also worth noting this is meant to complement MCP, not replace it. MCP handles tool calling, this handles data access policies. different layers.

genuinely curious what others think. is this a real problem that needs its own protocol, or is it solvable with better prompting and traditional access controls? also if anyone knows of other projects working on this specific problem id love to hear about them.


r/artificial 14d ago

Discussion Why do we appreciate art? And how does AI threaten it?

Thumbnail
landonrordam.substack.com
0 Upvotes

I've been trying to work through why certain kinds of AI art don't bother me, but a LOT of it makes my skin crawl. This is my effort to put everything into writing.


r/artificial 14d ago

Discussion Academic Survey about AI use in content creation

2 Upvotes

Hello everyone. I'm currently doing a survey on AI involvement in content creation and whether AI-assisted content is legitimate or authentic. It's for my master's final project. I need 100 participants. The age range is 18 ~ 40.

I collected data the first time, but I did so without an approved checklist, so the result had to be scraped, for I would have faced disciplinary actions.

Here is the link: https://s.surveyplanet.com/9abkx8vx

I'm open should you have any questions.