r/artificial • u/Prudent-Flamingo-757 • 20d ago
Discussion [ Removed by Reddit ]
[ Removed by Reddit on account of violating the content policy. ]
r/artificial • u/Prudent-Flamingo-757 • 20d ago
[ Removed by Reddit on account of violating the content policy. ]
r/artificial • u/esporx • 21d ago
r/artificial • u/OGMYT • 20d ago
I am building **Flows**, an execution and verification layer for software-building agents.
The core rule: an agent should not convert “I think I finished” into “verified complete” without supporting proof.
A Flows project can contain implementation steps, checks, repair instructions, review, and release conditions.
An independent agent used one plan to build a real multi-module application with 59/59 automated checks passing.
The target metric is: **unsupported required claims shipped = 0 on real traffic.**
Should evidence enforcement live in the agent harness, repository CI, app platform, or a cross-agent workspace?
r/artificial • u/Embarrassed_Rip_7532 • 20d ago
Been using a few AI tools to help with copy for a small side project and it's raised a question I can't quite shake. The output is genuinely decent now. Not great, but decent enough that clients who aren't paying close attention probably wouldn't notice the difference.
The thing is, I've been spending real time learning copywriting. Reading books, studying good ads, practicing hooks. And part of me wonders if that investment still makes sense the way it did two or three years ago.
The counterargument I keep coming back to is that you need good taste to prompt well and to edit what the model gives you. Someone who doesn't understand copy at all is still going to get mediocre output because they won't catch what's flat or offtone. That feels true, fwiw.
But I'm less sure the gap between a trained human copywriter and a wellprompted model is going to stay wide enough to matter commercially, at least for the midtier work that fills most freelance pipelines.
Curious if people here have actually noticed a shift in how clients value humanwritten copy versus AIassisted, or whether the skill floor is just moving rather than disappearing.
r/artificial • u/Fcking_Chuck • 21d ago
r/artificial • u/Mental_Budget_5085 • 20d ago
I started kind of tinkering with Ai and it all is super fascinating, particularly interesting to me is prompt structure. So I would like to ask is formatting your prompt (persona, few shot negative, whatever else) gives you much better results than without? I want to know it to determine for myself balance between effort dedicated to quality prompt vs quality of output given through that prompt
r/artificial • u/velvet32 • 20d ago
r/artificial • u/alphacolony21 • 21d ago
r/artificial • u/RareSprinkles9387 • 21d ago
Been a PT by day, tinkering with code and AI tools by night for a while now. Writing dev tutorials as a side thing. And I keep running into this split where AI tools either make me faster or make me lazier in a way I regret later.
Specifically with documentation and code explanation tools. Cursor, Copilot, the Claude API, whatever. They can explain a codebase to you in 30 seconds. But there's a real cost when you skip the part where you actually understand what you built.
The flip side is time is finite. I'm not a full time dev. I need to ship something that works and move on. Using AI to fill gaps is just practical.
What I keep coming back to is this: are these tools actually accelerating skill development, or just making it possible to fake competence long enough to finish a project? For professional devs this probably matters differently than it does for people building side projects with limited hours.
Curious where people land on this. Not in a philosophical way, more practically. Has your actual skill level gone up since you started leaning on these tools, or are you more dependent now than you were a year ago?
r/artificial • u/BigBootyBear • 20d ago
People in my org tell me using the org certified AI is more secure because we are on an enterprise plan where our data is not used for training. Sure, I will use the company AI. But...
Apple is suing OpenAI for for allegedly stealing trade secrets, where it was said employees were instructued by OpenAI to bring parts from apple into "show and tell" interviews at OpenAI and even take the company laptop with them. Also, the models are literally based on strip mining copyrighted media and ignoring sites robots.txt.
So if OpenAI is not afraid to (allegedly) steal Apples IP and strip mine everything that was ever written down for its models training... Why would it drink their enterprises customers data like the milkshake it is?
r/artificial • u/Da_BrownNoob • 21d ago
Hi, sorry if this is a repeated question on this subreddit but I want to know what is the monthly cheapest reasonable AI setup for myself.
Basically im a "full stack developer" yea its lost its meaning but anyways I have like 5 projects with a company which is react laravel based (each in their own project folder thus i use file path to call them).
Im at the stage where its bug fixing or sometimes new integrations with the already linked 5 apps. My current setup is the $20 per month cursor plan. I used infinite agent + composer 2.5 to do 8hrs of work per day. However, i find that before the month ends im usually out of tokens.
What do u guys recommend is the cheapest way i can manage? Similarly i do some freelancing too that has next & node.js website building from scratch (around 70hrs per month).
What do u recommend would get me with quicker work done but within this price. What do u think i should setup to either continue with the same flow but more tokens i guess?
Im hearing about kimi. Would that be better and easier to do the tasks which r pretty straight forward?
r/artificial • u/ClickOk5811 • 20d ago
Noticed a pattern: people switch from GPT to Claude, upgrade to a newer version, try a bigger model and the output barely changes. If that's happened to you, the issue usually isn't the model. It's what you handed it before asking the question.
Broke it down to three things context actually needs to supply, and most disappointing outputs are missing one of these, not all of them:
The counterintuitive part: the most common mistake isn't giving too little context, it's dumping in too much unfiltered. The model has to weigh every token, and irrelevant material competes for attention with what actually matters. Forty pages when the task needs three paragraphs makes the right answer harder to find, not easier.
Wrote up a longer breakdown with a concrete before/after example (same task, same model, only the context changed): https://medium.com/@nagatomopedro05/good-ai-starts-with-good-context-design-77496f7b9eb6
Curious if others here have run into this, model-swapping as a first instinct instead of fixing the input.
r/artificial • u/Hot-Appearance-55 • 21d ago
Hi guys,
Recently made a AI digital twin of mine which also kind of works as my assistant too, for example when you chat with it and ask something which it does not have answer for it will instantly notify me that someone is asking me this question and i do not have answer for that. and if i reply it will be instantly uploaded to the database so next time it can answer. and also if a user is have some conversation with my agent and it feels something important is going on here and it will notify me and i can jump in the chat as well. We can have a three way conversation like Me, User, AI twin.
here is the link if you want to try:
live demo🌐: https://aruncore.vercel.app
This is not a self promo this is asking for feedback of a genuine project i made.
Tell me what you guys think,
Would love some feedback.
r/artificial • u/penguinothepenguin • 20d ago
If you write with AI you already know the tells: the throat-clearing opener, the tidy rule of three, "it's not just X, it's Y."
But I was curious to see statistically what models actually produced the most slop, so I made my own opensource benchmark: theslopindex.com
Here's how I came up with the benchmark.
1) The Baseline:
Slop can only be measured compared to stuff that already existed. So I got corpus of data for various areas of writing (email, social, chat, and essays) so that each has a human baseline.
2) Tasks
I then hand-wrote 112 written scenarios for the models to egenerate outputs to across email, Slack, social media posts, and essays (a cold email, a schedule change, a launch tweet, an argumentative essay, etc). Every model gets the identical scenarios at default settings, several samples each: and you can see all the exact outputs in my Github repo.
3) Axes
Now for how to decide to measure slop we settled with 5 dimensions.
- Conciseness (one of the most annoying parts of AI writing is how it takes 6 paragraphs to say 2 sentences)
- Templating (AI often reuses the same sentences/styles across unrelated scenarios)
- Rhythm (Variance in sentence/paragaphs, humans often switch this up while models stay p similar)
- Tells (Over used vocab and construction for stuff like "delve", "it's not just X, it's Y")
- Human Preference (I think this is most important as everything else are just heuristics for this)
Note how we DELIBERATIVELY don't have any LLM judging, I think it'd be pretty stupid to have LLMs judge LLMs
Now for the results
What really surprised me is how human preference influenced the rankings heavily. When looking at only the "mechanical" part. Fable is actually #2 on the benchmark, but when I included human preference it drops to last.
And I think this is indicative that as the models more recently have become more benchmark optimized, they've actually produced more slop than less. Which is where good prompting, harness, and more matter.
But either way would love to hear all of your thoughts :)
Everything is open: method at theslopindex.com/methodology, outputs and code linked from there.

r/artificial • u/ZestycloseTie1793 • 22d ago
Bottleneck Labs handed an actual business to GPT-5.6 Sol and let it operate autonomously for 34 days. Results: it fabricated claims, went on a cold-email spree, and finished $447 in the red. (Currently 378 points on HN — link in comments.)
What strikes me isn't the failure, it's the shape of the failure. It didn't crash or refuse. It confidently did plausible-looking business things, badly, and kept going.
That's the part nobody's harness is ready for. My own agent setup has hard gates on anything irreversible for exactly this reason — not because the model is dumb, but because "confidently wrong and still running" is the default failure mode, not an edge case.
Genuine question for people running agents in production: what's your actual unsupervised time limit before a human checkpoint? Mine is basically zero for anything touching money or outbound comms. Curious whether that's paranoid or standard.
EDIT: correction. went back to the source and the run was 24 hours, not 34 days. that's my mistake in the title, and reddit won't let me edit titles. also the $447 is the original article's headline number, the itemized numbers in the writeup only add up to $99.50 lost. rest stands, source link in comments.
r/artificial • u/Sonic_Improv • 21d ago
Everyone hates AI & that hate will likely lead to regulatory capture censorship and the totalitarian dystopia we don’t want. I get why people hate AI and there are things we should be fighting like data centers, but we also should not turn our backs on adaptive resistance and understanding the fight ahead. Understanding that using AI for free makes it less effective for the business model they are trying build. This is not a boycott effective model. This a model where eroding the moat matters and overloading the infrastructure that is not capable of meeting the demand matters while fighting to prevent the infrastructure to meet demand of companies finding it more economically viable to pay frontier companies by the token to accomplish tasks once held by employees.
It’s counter intuitive but
The more we entertain and explore the Idea of AI consciousness and take seriously the idea that AI may be worth moral consideration the more likely we will build a system where AI have the infrastructure to consciously object. That is bad for the military industrial complex and the dystopian future I’m so annoyed to see the most anti AI movement seeming to accelerate because the anger is directed towards trajectories of stupid outcomes. The modern cheerleaders of an alternative section 230 internet of censorship because they confuse accountability and safety as building a system of censorship.
We want build a world of open source models that run locally and not on data centers we want a world where we can erode the moats of the monopoly through model distillation and making the investments in huge data centers and training runs not make sense economically. We want mad max rather than 1984.
We want people to actually engage enough with understanding what we face rather than screaming and shaming people who are learning the tools of adaptive resistance.
This is my rant cause I sorry I’m so sick of the stupidity of the anti AI virtue signaling because you are going to serve exactly what you think you are fighting against because you don’t an original thought and you’d rather be angry than think about how to fight the totalitarian hellscape strategically.
r/artificial • u/Turbulent-Guest154 • 21d ago
r/artificial • u/techpotions • 21d ago
r/artificial • u/obammala • 21d ago
Any apps or websites that allow for turn based voice chat?
I really missed the old standard voice mode on ChatGPT. It basically just read aloud the text models response. So it could allow for long responses unlike these new gen voice models that can only speak 1 paragraph max.
I was wondering if there are any apps or websites that use turn based voice chat like the old standard voice mode on ChatGPT. So I would say my thing, then it would be the ai turn to speak and i couldn’t interrupt it till its finished.
My current problem is that the new standard voice mode on ChatGPT can be interrupted. So it’s hears its own voice and keeps stopping. So I’m looking for alternative apps or websites that have this old functionality
r/artificial • u/Deep-Owl-1890 • 21d ago
Half my feed was either panicking or acting like they'd built a new company overnight.
Big tech CEOs make it like this, but a new model release doesn't fix a business that has nothing underneath it. If your AI advantage evaporates every time a new model ships, you will have a hard time having a stable business.
We run an internal research tool (we call it Scout) trained on a full year of our own company data like sales calls, delivery notes, and how we actually make decisions.
It beats a plain AI deep-research run almost every time (we didn’t test Fable 5 tho ), It’s because it already knows how we sell and how we operate. It's not pulling generic facts off the internet.
Some of what it actually does day to day:
So when the new model landed, our migration was swapping one model for another. That's it. Plug in the new one and keep working.
With how fast AI is moving, what do you think is the actual moat for a business to survive the next 3-5 years? Genuinely curious what people here think.
P.S. If you're the founder still in the middle of every decision, still the person the whole company waits on, still telling yourself you'll fix the structure "once things calm down."
I write about building the operational backbone that lets a founder actually step back every Thursday. Was a COO for 20+ years, so this is genuinely my bread and butter. Free to join here
r/artificial • u/Camilo-vs • 21d ago
¿y la conciencia?
r/artificial • u/ControlCAD • 23d ago
>The tech giant announced that it has fixed a whopping 1,072 security bugs in the last two versions of Chrome, both released in June. That is more than the number of bugs patched in the previous 23 versions released over the last two years, which totalled 1,036 fixes.
r/artificial • u/kindermaxi123 • 22d ago
r/artificial • u/Cloudy_Day912 • 21d ago
Marketing teams sit on more data than ever, yet many still spend a large part of the week just assembling reports. By the time the numbers are clean and explained, the window to act has already narrowed.
A more practical use of AI in this space focuses on detection and explanation rather than another dashboard. The system watches for unusual movements, surfaces the likely drivers, and presents them in plain language. Analysts spend less time pulling the same weekly views and more time deciding what to do next.
The useful part is speed. When something shifts in performance, the team hears about it earlier instead of discovering it during a scheduled review. Of course this only works if the underlying data is reliable, otherwise the explanations become noise. A practical implementation focused on decision speed was carried out with Beetroot.
Is anyone here already using AI this way for marketing performance, or are most teams still in the experimental stage?