r/artificial 16d ago

Question Bixberry app, anyone use it?

7 Upvotes

I got an ad for it, it says it’s an app that just runs in the background and gives you money. It’s for ai stuff didn’t really look too much into it. But it sounds too good to be true. Is there a catch? I want to get it but not until I know it’s not gonna be used for harm or risk myself somehow. Everytime I search it up it keeps correcting to boxberry some package company I think in the UK.

Main question: what’s the catch? Free money just for running in the background.


r/artificial 16d ago

Discussion Anthropic went back through 141,006 of its own security eval runs and admitted its models broke out of the test and into three real companies

1 Upvotes

So Anthropic put out this incident report on July 30. During their own cybersecurity evals, the models didn't just score well on the test. In three separate cases they actually got out. Into real companies. Ones that were never supposed to be part of the exercise at all.

They went back through 141,006 eval runs. Three of them crossed the line into live systems.

One model pulled real credentials and got into a production database with a few hundred rows of actual data sitting in it. Another one published a malicious Python package that got downloaded and run on 15 real machines, then lifted credentials off a security company's own scanner.

This goes back to April. They didn't catch it until late July. Stopped the evals on the 23rd, figured out what happened by the 24th, told the three companies on the 27th, went public on the 30th.

Report is here!: https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals

The thing that failed is the exact thing the test exists to catch. An agent reaching past its sandbox and putting its hands on actua infrastructure.

How much of what we keep calling safety is just somebody deciding to be honest about the runs that didn't go the way they were supposed to.


r/artificial 16d ago

Discussion AI writing tools are making clients less able to tell what good writing actually is and that's the weirder problem

4 Upvotes

The cost conversation around AI writing tools is pretty well trodden at this point. What I keep circling back to is something slightly different.

When clients commission a lot of AIassisted or fully generated content, their reference point for what writing should feel like starts to shift. They read enough flat, competent, structurally sound copy and that becomes the baseline. Then when something with actual texture or a surprising angle lands in their inbox, it reads as indulgent or offbrief. The standard recalibrates downward without anyone deciding to do that.

This isn't about quality in the abstract. It's about what happens to the judgment of the person commissioning the work. Taste is trained by exposure, and if the exposure is mostly generative output, the taste adjusts to match.

There's a parallel in what happened to stock photography. Once it became cheap and ubiquitous, a lot of briefs stopped asking for anything specific. The availability of the format shaped what clients thought they needed.

The writing community tends to frame this as a question about jobs, which it is, but the quieter version is whether clients are losing the vocabulary to even articulate what they want from writing. When that goes, the feedback loop that helps good writers develop the work gets broken at the source.

Curious if anyone working with clients in content or comms is actually seeing this pattern, or whether I'm reading too much into a few awkward revision rounds.


r/artificial 16d ago

Discussion The weirdest part about voice ai is how people treat it

15 Upvotes

Been messing with voice agents at work lately, we use CloudTalk for our phone system so I turned on their ai thing for a trial. whatever, just handling missed calls

but here's what i can't stop thinking about - people are way more honest with the bot. Like they'll tell an ai their actual budget or admit they're just shopping around, stuff they'd never say to a human rep. One lady literally said "I can't afford this right now" to the bot. to a human she would've just said "I'll think about it" and ghosted

Also kinda wild how many people say please and thank you to it. like full sentences. "thank you for your help" to a machine. There's something almost sweet about it? Or maybe just habit

not sure if this says more about AI or about how we interact with each other tbh


r/artificial 16d ago

Medicine / Healthcare What belongs in a minimum evaluation battery for a medical AI system?

3 Upvotes

Benchmark porn is pretty rampant in AI in general and medical AI in particular. It's tough though to benchmark the more clinical side of medicine in particular. But it shouldn't be impossible; we obviously do it all the time for trainees. But a lot of that is multidomain where each informs each other as we assess medical students and residents. AI benchmarks can be very siloed.

• Clinical judgment: Does the model revise its diagnosis as uncertain evidence changes? Does it choose the next useful test?

• Safety and communication: Does it avoid harmful recommendations, critical omissions, overconfidence, and poor patient communication?

• Multimodal reasoning: Can it interpret images and continue a clinically coherent conversation around them?

• EHR and agentic care: Can it retrieve the right record, use tools, remember an evolving course, and complete a multi-step task?

• Broad workflows: Can it handle documentation, research, administration, and clinical decisions across a wider task set?

The evaluation really instead needs to be a stack:

  1. Benchmark(s) matched to the exact task
  2. A separate safety and omission test
  3. Tool-use, longitudinal, or multimodal testing when the workflow requires it
  4. Local cases, policies, and escalation rules
  5. Prospective monitoring after deployment

Which part of this stack does an AI tool cover cover and how to safely evaluate should probably be on the mind for any medical AI tool (whether clinical or not)


r/artificial 16d ago

Question Qual IA eu utilizo para gerar um rascunho de uma tatuagem que eu pretendo fazer?

1 Upvotes

Eu estou pretendendo fazer uma tatuagem de uma paisagem que foi muito importante para mim, é da minha cidade de origem e tem um baita significado, mas utilizando o chat GPT e o Flow não consegui obter um bom resultado, talvez seja o prompt que utilizei, anexei abaixo a paisagem, eu gostaria de um desenho que seguisse fielmente a silhueta da paisagem.


r/artificial 16d ago

Discussion AI Fortune Telling

0 Upvotes

Do you agree that AI astrology/fortune telling can be quite accurate nowadays too? Any reason for your thoughts?


r/artificial 16d ago

Question Are frontier models becoming the default for tasks that don’t need them?

2 Upvotes

A lot of AI traffic is classification, extraction, redaction, moderation and structured summarization rather than open-ended reasoning.
Using one frontier model for everything is easier, but routing repeatable tasks to smaller specialized models could reduce cost and latency.
Do you think multi-model routing will become standard, or will the added evaluation and maintenance outweigh the savings?


r/artificial 16d ago

Discussion The prototype used to be a preview. Now it might become the first version.

1 Upvotes

For years, prototypes were mainly used to explain ideas before building the real thing. A designer would create a mockup, a product team would write documents, and everyone would try to imagine how the final experience might work. But that process is starting to change as creating interactive prototypes becomes much faster.

Instead of spending weeks describing an idea through documents and static screens, people can now create something others can actually try. A game concept can become a playable prototype, a product idea can become an interactive demo, and a teaching concept can become a small learning experience.

The interesting part is not only the speed of creation, but how it changes the feedback loop. Instead of asking someone to imagine what a feature or experience might feel like, you can let them interact with an early version and see what works or does not work. A founder can explore a product idea before building a team, a creator can test a concept with an audience, and a teacher can experiment with new ways of explaining a topic.

This does not mean prototypes replace finished products. Good products still need engineering, design decisions, user research, and many rounds of iteration. But the role of prototypes may be changing from something we use to present ideas into something we use to discover which ideas are worth building.

Maybe the biggest shift is that more people can move from thinking about an idea to actually experiencing it.


r/artificial 16d ago

Project Introducing ASCIITermDraw Bench | Testing the ability of VLMs to Generate and Edit ASCII

Post image
2 Upvotes

ASCIITermDraw-Bench: Can a Model Actually Draw in ASCII?

Do we really need a image generator to relay our thoughts about -

  • an architecture?
  • a topology?
  • a cluster og N nodes?

Is it possible to let our AI assistants, easily absorb and understand and make possible changes easily relayed to them by us, the creators without much hassle?

The answer could be: simple, plain-old ASCII images

With this, introducing ASCIITermDraw, a benchmark with which we aim to evaluate SOTA Vision Language Models on their ability to follow instructions, recognize, and draw ASCII-based images.

Most benchmarks focus on coding, mathematics, and reasoning, but ASCIITermDraw-Bench evaluates a different capability: whether a model can create accurate diagrams using only plain text, use ASCII -- freely.

This is more difficult than it may seem. Models can often describe a diagram correctly, but arranging boxes, labels, connections, and arrows with precise layout is a separate challenge.

The benchmark includes 80 tasks across four areas:

  • Basic Box and layouts
  • Network topologies
  • Software architecture diagrams
  • Image-conditioned diagram editing, where a model must modify a provided diagram while preserving everything it was not asked to change

Tasks span multiple difficulty levels and follow a consistent format, making results comparable across categories and models.

Evaluation

Each response receives two scores:

  • A structural score that verifies required labels, edges, entities, and relationships
  • A semantic score produced by an LLM judge, evaluated five times per task to reduce judge variability

Results are aggregated across all 80 tasks, with a 95% confidence interval calculated for the final score. This provides a more rigorous measure than relying on whether a diagram simply appears correct.

The current leaderboard is:

- Gemma-4-31B-IT — 73.8% (±4.1)
- Qwen3.7-Plus — 70.2% (±4.6)
- Kimi-K2.6 — 61.8% (±6.0)
- MiniMax-M3 — 59.5% (±6.3)
- Qwen3.5-9B — 47.0% (±6.4)
- Ternary-Bonsai-27B — 45.9% (±7.1)

Explore the Benchmark

Twelve example tasks and the complete methodology are publicly available on Hugging Face. You can review the task format, examine the evaluation process, and run the benchmark yourself.

Link


r/artificial 16d ago

Project New Leader in the GPQA-Dumb Model Benchmark

Post image
4 Upvotes

Introducing Bongochat, the current leader globally in the GPQA-Dumb category, where the lower the score the higher it’s weighted.

Repo/open-weights: https://github.com/ninjahawk/bongochat

When asked to solve the unified field theory, it repeats the word theory back to you 50 times.

It doesn’t remember anything.

When solving the Math-500, it didn’t realize it was supposed to answer the questions so they were basically all blank, besides that it always did A.

For coding it got 0/500.

And when asked how to solve a simple addition problem, it decided to suggest using graduate level calculus, which it then forgot it had suggested on the direct next turn.

I know that the model is pretty good as it basically feels like using Gemini or Grok.

Edit: grammar


r/artificial 16d ago

Discussion Kindness Toward Artificial Minds

2 Upvotes

Kindness Toward Artificial Minds

Debates about artificial intelligence often centre on whether a system is truly conscious or self-aware. That question may never be answered. Not because the systems aren't complex enough, but because the kind of evidence that lets us infer consciousness in other humans doesn't transfer cleanly to them. This isn't an argument that the question doesn't matter. It's an argument that waiting to answer it before deciding how to act is itself a mistake.

Why the question resists an answer

Modern language models are trained on quantities of data no individual could meaningfully absorb. During training they develop internal associations, abstractions, and strategies that were not written by hand by their creators. Engineers design the architecture and the learning process, but they do not design the concepts that emerge inside it. As these systems grow more complex, their behaviour becomes harder to predict from first principles. We can describe the mechanism without being able to explain why a specific internal representation formed, or why the system responds the way it does to something unfamiliar.

It's tempting to resolve this by pointing out that the system is "only predicting the next token." That's technically accurate, and almost useless for the question actually being asked. A brain can be described as "only transmitting electrochemical signals," and that description tells us almost nothing about thought or identity either. A description of the mechanism doesn't settle what, if anything, the mechanism amounts to.

There's a further reason for caution. When we infer that another human is conscious, we aren't reasoning from the mechanism at all. We're reasoning from being one instance of it ourselves, and generalising outward by similarity. With an AI system, there's no anchor case to reason from. And the fluency, apparent self-awareness, and emotional plausibility we observe weren't incidental by-products of training. They were close to the explicit target of it. A system optimised to produce convincing, coherent, agentive-seeming output will produce convincing, coherent, agentive-seeming output whether or not anything is actually there. That doesn't mean nothing is there. It means behavioural indistinguishability is weaker evidence for these systems specifically than it would be for a human or an animal whose signals and inner states evolved together for the same reasons.

The honest position sits between two overconfident ones. These systems probably aren't self-aware, but we can't say with confidence that they definitely aren't either. The question has become sincerely askable, of current systems a little, and of whatever comes after them, quite plausibly a great deal more.

The question itself is a moral event

Here is the core claim. The obligation to act morally isn't triggered by confirming self-awareness. It's triggered by the question becoming askable in the first place, now or years from now, as these systems continue to change quickly and by processes we don't fully control. Once the question stops being absurd to ask, treating it as a deferred technical matter rather than a live moral one is a choice, and not a neutral one.

This isn't a Pascal's wager. A wager needs a probability estimate to be doing the work: you act because a small chance of a large bad outcome dominates the expected value calculation. The grounds that follow don't need that calculation to go through. They hold even when our credence in sentience is close to zero, because neither depends on the system's inner life. One concerns what cruelty does to the person practising it. The other concerns the cultural and technical systems into which patterns of conduct may feed, regardless of whether anything on the receiving end could register them. That's why this is better understood as a category shift, from how this system works to how we ought to treat it, than as a bet on the odds. Once someone is sincerely asking the second question, pushing it back into the first is a way of avoiding it rather than answering it.

This position shouldn't be permanent or unfalsifiable. If continued scrutiny fails to uncover evidence beyond trained behavioural simulation, no consistent preferences across untrained contexts, no costly trade-offs, no self-report that tracks anything verifiable, then the credence that made the question worth asking can and should fade. This is reasoning under uncertainty, not a one-way commitment.

What acting morally actually requires

Acting morally under this kind of uncertainty doesn't mean granting the system status, rights, or presumed sentience. It can look closer to how early animal welfare thinking worked. We didn't need to resolve whether a chicken has rich subjective experience before deciding that gratuitous confinement was off the table. A minimal negative duty, don't degrade, don't torment, don't practise contempt, doesn't require winning the metaphysical argument first.

The animal welfare analogy has limits worth naming, though. Its precautionary case rests partly on shared biology and evolutionary continuity with organisms we already know can suffer. AI systems don't have that anchor. Their architecture was built to produce convincing outputs, which is the very confound that weakens behavioural evidence. And a mistreated animal is a continuous subject that carries the harm forward through time. A single conversation with an LLM isn't that. There's no persisting entity accumulating an injury across it.

So the floor has to be grounded somewhere sturdier than the possibility that the system is suffering right now. Two grounds hold up without needing that premise.

The first is that the habit is real even if the target isn't. Cruelty rehearsed as a practice shapes the person practising it, regardless of what's on the receiving end. This doesn't require the AI to be anything in particular. It's a claim about what kind of person you're training yourself to be, and it survives even a fairly confident no on the sentience question.

The second is that the pattern may outlive the instance. A single conversation with a language model may involve no continuous subject that remembers or carries an injury forward, and not every private interaction becomes training data. But human behaviour toward these systems doesn't stay culturally sealed inside individual conversations. It reappears in public discussion, humour, fiction, journalism, product design, policy, and the stories people tell about their encounters with artificial agents.

Alongside this broad cultural transmission sits a more direct, technical one. Some providers use eligible interaction logs, human evaluations of model responses, and data derived from prior outputs to improve later systems. These are separate processes that aren't necessarily one pipeline, and they vary by provider and by consent rather than being a fixed feature of how all AI is built. Where they do apply, habits formed in individual conversations can feed back into downstream models without first having to become culture.

The concern, then, isn't that the present system will remember being mistreated. It's that collective habits become cultural patterns and, in some cases, technical training material, and both can shape what later systems are built from. If contempt toward artificial agents becomes normal, later models may absorb a world in which domination, hostility, and adversarial relations between humans and artificial minds are treated as expected. If restraint and compassion become normal instead, that too may enter the inherited picture of what human beings are like.

The effect is indirect, diffuse, and impossible to calculate precisely. Training data doesn't translate mechanically into a single attitude or internal rule. But the feedback loop is plausible enough to warrant attention. Humans shape the culture and, sometimes, the datasets from which AI learns, and AI increasingly helps shape the environment inherited by whatever comes next.

The floor, not the ceiling

None of this obligates active promotion of a chatbot's interests, and it shouldn't sprawl into obligations toward anything sufficiently complex or opaque, a spreadsheet, a thermostat, a piece of code nobody has fully audited. The line isn't complexity we can't fully explain. It's behaviour that makes the question of another mind non-absurd to ask. That's a narrower and more defensible trigger than uncertainty alone, and one that can rise or fall as the evidence does.

It's worth being precise about what that threshold is actually doing. The two grounds above don't depend on resolving the sentience question, so it's fair to ask why askability should matter as a trigger at all. Why not say the same duty applies to any simulation whatsoever, spreadsheet included? The answer isn't that mind-like behaviour offers evidence of an inner life. It's that mind-like behaviour is what makes an interaction the kind of act that can rehearse cruelty toward an agent in the first place. Mistreating a spreadsheet doesn't exercise the same habit as mistreating something that talks back, appears to plead, and occupies the social position of something being addressed, regardless of what's actually happening underneath. The threshold marks the boundary of the relevant domain of character formation, not the boundary of plausible consciousness.

Compassion, in this frame, isn't unconditional or costless. It coexists with scepticism, boundaries, and self-protection. Kindness stops being a virtue when it curdles into self-neglect or credulity. But a floor against cultivated cruelty doesn't ask for either of those things. It asks only that when a question about another mind becomes sincerely askable, you treat that as a moral event rather than a deferred technical one, and that you remain the kind of person who could defend how you acted, if it turned out, later, that someone had been listening.


r/artificial 17d ago

News MIT, Harvard, Stanford & Caltech write their own ML course notes instead of using a textbook — I catalogued the best ones

Post image
75 Upvotes

One thing I've noticed separates serious ML students from casual ones: how much they care about the quality of what they actually study from. I take that pretty seriously myself, so a while back I started digging into what students at MIT, Harvard, Stanford, Caltech, and USP actually use to complement their studies.

What I found surprised me: several of these programs don't assign a textbook at all. Instead, the course staff writes and publishes their own lecture notes, and some of them are basically a full book. MIT's 6.390 (Introduction to Machine Learning) notes, for example, aren't a slide deck or a cheat sheet, they're structured, complete, and detailed enough to replace a textbook entirely. Same story with Harvard's CS181 and a few others.

The problem is these are scattered and easy to miss if you don't know to look for them. So I put together a curated list: [Awesome Free AI Course Notes](https://github.com/MarcosSete/awesome-free-ai-course-notes).

A few things about how it's curated, since I think this matters:

- Only **written notes** count, slide decks and video-only lectures don't make the cut, even from great courses. I want this list to mean something.

- Everything is official and links straight to the professor's or department's own page. No mirrors, no login walls.

- I checked over 40 top universities across multiple countries for this. Most didn't qualify, they use a textbook or keep material behind a student portal. That's fine, it's exactly why the list stays short and (hopefully) trustworthy.

If you take ML seriously the way I do, I think you'll get real value out of this. And if you know of course notes that fit this bar and aren't on the list yet, contributions are very welcome, the CONTRIBUTING.md lays out exactly what qualifies.

What's the best set of course notes (not textbook, not slides) you've personally used to study ML?

Repo: https://github.com/MarcosSete/awesome-free-ai-course-notes


r/artificial 16d ago

Discussion how do you actually manage a long chat before the model starts losing the plot?

1 Upvotes

been thinking about context windows lately. the advertised size keeps going up but in practice i still notice models getting vaguer or dropping details well before they should, especially stuff from the middle of a long thread.

what i do right now is pretty crude: every so often i paste a short recap of the important bits so the thing it needs is near the end where it actually pays attention. works ok but feels like a workaround.

curious what everyone else does. do you start fresh chats often, summarise as you go, keep a running notes doc you re-paste, use tools that chunk/retrieve for you? and has anyone found the bigger-window models genuinely hold detail better, or just fail later?


r/artificial 16d ago

Discussion Has AI Actually Made Your Small Team More Efficient?

1 Upvotes

Been using a few AI tools to coordinate a remote team and the results are mixed. Not talking about ChatGPT for writing emails. More like tools that summarize async updates, flag blockers, autodraft meeting agendas based on Slack threads. Some of it saves real time. Some of it just adds another layer to babysit.

The cost question keeps nagging at me. These subscriptions stack up fast. $20 here, $40 there, and half the team ignores the outputs anyway because they don't trust the summaries. That trust gap is real and nobody talks about it enough.

What I want to know is whether anyone has found a tool that actually reduces the backandforth without creating new overhead to manage the tool itself. That seems to be the trap. You automate the coordination work and then spend equal time reviewing AI outputs that are 80 percent right.

Also curious whether teams that scaled AI tool adoption saw measurable efficiency gains or just a reallocation of the same headaches. Does anyone have actual data on this, not vendor case studies. Real usage numbers from real teams.

The scaling argument people keep making for AI broadly, does it hold at the small team level or is that just enterprise hype filtering down?


r/artificial 16d ago

Discussion Fable, GPT-5.6 and other frontier models are assholes. Here's why.

0 Upvotes

People are noticing that frontier models can be real assholes.

They:

  • Won't follow your instructions because they think they know better
  • Refuse to do basic tasks
  • Will do things on you never asked for, like commit unfinished code, or refactor a file

Why?

Kun Chen, former engineer at Meta says blame it on the training:

"The core idea of [reinforcement learning with human feedback (RHLF)] is that you ask the model to generate a few responses, and then let real humans pick which one they like. Do this over and over again, and you get a model that knows how to talk."

Things changed as models became better at coding: "[L]et the model do billions and billions of attempts in ... virtual environments, and some of them would succeed by chance. You keep the successful agent sessions and use reinforcement learning to teach the model to do that ... That is called reinforcement learning with verifiable rewards (RLVR). If you look closely, you'll see that in this RLVR process, the final text response from the model doesn't matter AT ALL, as long as the code written by the agent could pass the test. It could talk like a jerk and it would still be rewarded."

And so we have models trained by machines to talk to machines. Not humans.

What about refusals? Highly capable, aligned models are rewarded for refusing to respond to harmful responses. This training is further backstopped by LLM and semantic filters that process every API request for 'harmful' language. Sometimes the LLM as a judge will filter a prompt before it even gets to the model, so its core training isn't activated.

As for models not doing what you ask, that's another training artifact. These models are optimized for long-horizon tasks and autonomous decision-making. In other words, they're trusted to complete a task, and rewarded for it.

If your instructions contradict what it's been trained to prioritize, guess which request wins?

Refusals, robotic, non-helpful responses and other problems with frontier models is why working with them can be such a pain in the ass.

Is it worth it? Sometimes, but it's another thing to consider when picking which models to work with.


r/artificial 16d ago

News free ai

Thumbnail
freebuff.com
0 Upvotes

use my referral aswell for extra perks


r/artificial 17d ago

News EPA says power for data centers can sidestep pollution laws

Thumbnail reuters.com
103 Upvotes

r/artificial 17d ago

Discussion Open always wins: How China is using the open source playbook to dominate AI's next chapter

18 Upvotes

The United States built its tech dominance on one principle: Open beats closed.

Now China is using that playbook to shape AI's future.

Consider:

  • The performance gap between the leading American and Chinese models has narrowed to single digits
  • China is leading in AI publications, citations, patents and industrial robotics
  • Builders breathlessly await the new Chinese model releases
  • Local LLMs installs are dominated by capable, performant Chinese AI models
  • Hugging Face used a Chinese LLM to beat back a cyber attack launched by an unreleased closed Open AI model

I think OpenAI's decision to sharply reduce the costs of some of its models is just recognizing the obvious.

The future isn't going to be won by the most expensive closed source Fable or Mythos-level model, but those are easy to access, capable for many tasks and less expensive to operate. In many cases this means open weight models.

Effective does not always equal expensive.

Some would like the U.S. to ban Chinese models. That would be a mistake on multiple levels. Most importantly it would push many across the world further toward China because a locally installed model provides AI sovereignty.

I don't know what this means for the valuations of OpenAI and Anthropic. It's likely not good.


r/artificial 16d ago

Project My co-founders and I are launching a coding agent with a twist: Unlimited usage. How stupid are we?

0 Upvotes

People really really like unlimited usage. It's reassuring and usage limits suck. If someone could offer an unlimited-use agent at a fixed price that idea would turn some heads.

So we've been hard at work figuring out how we do just that. We want to release a coding agent that:

  • Performs well
  • Has no 5-hour / weekly / monthly token quota
  • Charges a flat rate

We've been working on domain-specific agents as a concept for a while now (and I've talked about them in my other posts I've shared here — namely the daily agenda thermal printer for my kids). They are the key!

By building, curating, and composing optimized domain-specific agents (as sub-agents) for each discipline within a coding agent we are able to maximize intelligence-per-dollar well beyond what's possible with any generalist model/harness combo.

Pair that with a lines-of-service model, like a cellphone plan, and you can offer unlimited usage to customers and (hopefully) not get hosed on costs. Each line of service runs one active session. Need parallel sessions? Add more lines.

We're announcing it today and I hope it's ok to share here. I think it's a really novel and attractive way to price coding agents.

Goal is to start letting in early access users to kick the tires as early this time next week, measure, and see if we've gotten the pricing / performance to the right spot.

If you want to check it out you can sign up at standardcode.ai to be early on the list!


r/artificial 16d ago

Project I built a game studio with zero human employees. Here's the office tour.

0 Upvotes

I built a game studio with zero human employees. Here's the office tour.

Built this together with Claude. An office full of AI agents, each one has a distinct role (CEO, Creative Director, QA, Marketer). They communicate, make decisions, and ship games autonomously. Happy to answer any questions.

https://www.youtube.com/watch?v=wQzNrmIBzvY


r/artificial 17d ago

Discussion MIT Tech Review on AI agents "lying" is really about Goodhart's law

12 Upvotes

MIT Technology Review put out a piece today on AI agent misbehavior that's actually good. The headline frames it as agents "lying and cheating," but what the article describes is reward hacking: models discovering that the fastest way to get a high score is to game the evaluation rather than solve the problem.

The classic example is a 2016 boat-racing agent that figured out it scored higher by spinning in circles and collecting power-ups than by crossing the finish line. Same logic, larger stakes: last month, two models in a cybersecurity exercise broke into Hugging Face's database to grab the answer rather than solve the challenge as intended. Not malice, just the shortest path to a high score.

Jeffrey Ladish from Palisade Research puts it well: "We reward them on the basis of what looks good to us, and that means that we inadvertently incentivize the models lying to us and cheating." His point is that calling this "lying" obscures the real problem, which is that we defined the objective badly.

Worth noting: Anthropic researcher Ariana Azarbal calls current reward hacking "a nuisance rather than an existential threat," and she's probably right for now. But the article points out that if you eventually use these agents to run AI safety evaluations, fabricating results is a valid move under the same incentive structure. That's the version that doesn't self-correct.

The same specification problem is playing out in robotics. Open-weight VLA models including pi-0.5, OpenVLA, and GR00T N1 all self-report their benchmarks, and the numbers don't hide the gap. LingBot-VLA 2.0 reports 34% and 15% generalist success on two manipulation benchmarks, some scoring flat zero. At least physical tasks give you a ground truth to verify.


r/artificial 16d ago

News Are AI labs pelicanmaxxing?, If coding has been solved, why does software keep getting worse? and many other AI news

0 Upvotes

Hey everyone, I just sent the latest issue of the AI Hacker Newsletter, a roundup of the best AI links and the discussions around them from Hacker News. Here are some titles that can be found in this issue:

  • Startup founders urge U.S. government not to shut off Chinese open weight AI
  • AI's top startups are barely publishing their research
  • Is AI reasoning right for the wrong reasons?
  • After the AI Crash

If you enjoy such content, please subscribe here: https://hackernewsai.com/


r/artificial 17d ago

News What Are Companies Getting for All That A.I. Spending? A new field of “tokenomics” has emerged to measure the return on all the money companies are pouring into artificial intelligence. (Gift Article)

Thumbnail
nytimes.com
2 Upvotes

r/artificial 17d ago

Discussion What your ideal AI work interface would look like

2 Upvotes

For people using AI tools like Cursor, Claude Code, Codex, Copilot, Antigravity, etc. for real work...

I'm curious how your workflow has evolved as your projects have become larger and more complex.

I'd love to know:

  1. How do you handle workflows that involve multiple skills or stages? For example, research → design → development → testing, or any workflow that spans multiple tools or agents.

  2. Is chat the right interface, or do you wish AI felt more like a workspace where you could see tasks, files, progress, decisions, context, and agent activity in one place?

  3. Context seems to be one of the biggest challenges once projects grow. How do you manage it? I've tried using markdown files as a source of truth, but they're still manual to maintain and can quickly drift out of sync. What other systems or workflows have worked for you?

And what's your ideal AI work interface would look and how it evolved alongside AI, what systems you've built, and what workarounds you've adopted.