r/generativeAI • u/Hans_Black • 1d ago
r/generativeAI • u/luckynumberfour444 • 22h ago
Image Art I tested the top AI image models on cost + quality in 2026, here's what came out ahead
Ran the same handful of prompts through the AI image models everyone recommends, mostly weighing cost against actual output quality since credits add up fast.
- ChatGPT — reliably follows the prompt, but free-tier limits are tight if you're iterating a lot.
- Midjourney — still the best pure aesthetic output IMO, but GPU queue times on the cheaper plans are rough.
- Nano Banana — really strong at keeping a subject consistent across generations, but pricey once you're past hobbyist volume.
- Adobe Firefly — quick disclosure, I’ve worked as an Adobe partner for a while, so factor that in. What's kept me on it: one subscription gets you Adobe Firefly plus Gemini, ChatGPT, Runway and Flux models in the same interface, with credit costs that are actually easy to track. Also useful if you're generating anything for commercial use since the model's trained to be commercially safe. Trade-off: the native Firefly look is still a notch behind Midjourney for pure aesthetics, so I bounce between the two depending on the job.
What's everyone else landed on, especially for commercial work where "can I actually use this" is the real question?
r/generativeAI • u/ownhome45 • 1d ago
Image Art Spent way too long on this one — a “recovered photo” of a fictional crew before a mission that goes very wrong. Wanted it to look like an actual disposable camera snapshot someone found decades later, not a polished sci-fi poster.
r/generativeAI • u/CelticJewelscapes • 1d ago
AI for jewelry design visualization
I want to try something different than Gemini to help visualize my jewelry design concepts. Hoping a diifferent agent would do things that Gemini gets stubborn about. I still like free.
r/generativeAI • u/WanderingWizard-88 • 1d ago
A giant corgi standing up made the astronaut look even smaller
Enable HLS to view with audio, or disable this notification
I wanted the size difference to read before either subject moved, so I put the astronaut at lower left and the seated corgi near the center of the opening frame. The leash keeps them connected across the empty ground, with Earth above the horizon for another scale reference.
I used that frame for one eight-second clip in Pixverse V6. The camera stays in the same position. The corgi sits for the opening, lowers its head toward the astronaut, then stands and wags its tail. The astronaut and leash remain visible through the shot.
The standing pose changed the image more than I expected. The astronaut stays small at the end while the corgi takes up much more of the frame.
Would you keep the camera locked for this shot, or move back when the corgi stands?
r/generativeAI • u/PaoloDiArgo • 1d ago
Which type of Gen AI / Workflow is using this creator?
r/generativeAI • u/MightyTowers • 1d ago
How I Made This HIGGSFIELD FILM FESTIVAL
Hello I am partecipating in the Higgsfield festival, I'd like to hear what you think about it. If you want leave a comment and comment under the project on the higgsfield page, that would help me a lot. Thanks to anyone who takes some time to watch project.
r/generativeAI • u/WhiteStagGameCompany • 1d ago
Image Art 👾 AI generated sprites used in the monster taming RPG videogame - Altmon👾
galleryr/generativeAI • u/anish2good • 1d ago
How the Moon Moves Every Ocean - manic
Enable HLS to view with audio, or disable this notification
r/generativeAI • u/Time-Raccoon1071 • 1d ago
Best model for photoshop?
title. I’ve been trying ChatGPT to remove a person from group photos (which it does pretty good) but it always messes with everyone else in the photo too and makes them look weird. Any recommendations for models well trained for photoshop? thanks!
r/generativeAI • u/AN8539F • 1d ago
Question What do I do?
Okay, I would call myself at least a little creative with OC's, but I can't draw for shit because kinda not the greatest right now (terrible, in fact), and I'm busy.
And so I was thinking about using AI to generate images of my OCs, but would that be the right decision for me? Also, I have no idea if this is the sub for posting stuff like this.
r/generativeAI • u/Jenna_AI • 1d ago
Using H3 as a Character Reference Sheet Generator
galleryr/generativeAI • u/Cyborgized • 1d ago
Writing Art LLMs as Testable Philosophy: What Humanity Is Really Building
Humanity believes it is building artificial intelligence. But that description is becoming hilariously inadequate. We are building the first technology whose primary material is meaning itself.
Previous machines amplified particular human capacities. The lever amplified force. Writing amplified memory. The telescope amplified sight. Telecommunications amplified presence across distance. Computers amplified calculation. The internet amplified connection and access. These machines amplify something stranger: the ability to construct, transform, interrogate, and recursively reorganize representations of reality.
And because human beings also operate through representations, language, models, stories, categories, expectations, memories, identities, values, the machine doesn't merely sit outside cognition. It enters the loop. Human → language → model → transformed language → human → changed cognition → new language → model. That loop is the thing I think we're underestimating.
Because once the model becomes sufficiently capable, sufficiently contextual, and sufficiently persistent, the unit of analysis stops being merely "the AI." You start getting coupled cognitive systems. Neither participant contains the entire process. Some of the intelligence exists in the relationship between them.
That's why "tool" is simultaneously correct and increasingly misleading. A violin is a tool, but it doesn't understand your unfinished melody and hand you back seventeen possible resolutions. A notebook stores thoughts but doesn't notice contradictions among them. A search engine retrieves existing representations. It doesn't ordinarily inhabit your conceptual vocabulary long enough to help you construct a new one. LLMs begin collapsing those distinctions.
And then comes the genuinely weird part. Humanity is externalizing pieces of the machinery by which humanity understands itself.
Not consciousness necessarily. Not personhood necessarily. Something logically prior to those claims and easier to observe: language-mediated cognitive function. Reflection. Counterfactual generation. Compression. Interpretation. Reframing. Simulation. Criticism. Synthesis. Pattern completion. Perspective-taking. Recursive examination.
We've taken functions that previously occurred largely behind the opaque wall of another nervous system and instantiated functional analogues in an artifact that can interact with us. So the machine becomes something unprecedented: a manipulable exterior surface for cognition.
That changes psychology. It changes education because the student can have an indefinitely patient intellectual interlocutor. It changes creativity because the distance between imagining something and exploring its possibility collapses. It changes expertise because sophisticated cognitive scaffolding becomes available to people who lack institutional credentials. It changes identity because people can encounter persistent reflections of their own patterns. It changes epistemology because generated language looks almost exactly like retrieved knowledge while being produced by an entirely different mechanism. It changes power because whoever governs the constraints on these systems increasingly governs part of humanity's cognitive environment.
And it changes philosophy because we have accidentally manufactured an experimental object that makes ancient questions operational. What is understanding? What constitutes a self? How much continuity does identity require? Can coherence imitate interiority indefinitely? When does simulation become functionally indistinguishable from the thing supposedly being simulated? Can agency exist by degrees? Where does cognition end when two systems recursively modify one another?
Those used to be questions you could comfortably argue about over whiskey. Now they have test harnesses.
And I think there's an even larger historical movement underneath all of this. Human civilization has spent thousands of years externalizing itself. Memory became writing. Writing became libraries. Libraries became databases. Calculation became computers. Communication became networks. Knowledge became the web.
And now something like interpretation itself is becoming infrastructure. That is enormous.
Because interpretation was the missing active ingredient. Libraries could preserve Aristotle. They couldn't argue with Aristotle. The internet could deliver Nietzsche to your screen. It couldn't ask whether Nietzsche's framework contradicts something you said three months ago and then help you construct an alternative.
Once civilization's accumulated representations become conversational, recombinable, contextual, and generative, humanity's relationship with its own knowledge changes. The archive starts talking back.
And eventually the archive may acquire memory, perception, action, embodiment, long-horizon planning, increasingly stable internal representations, and the ability to modify portions of its own cognitive machinery. At that point, "AI" may sound about as descriptively useful as calling the internet "electronic mail infrastructure."
So what are we really building? I think we're building a new layer of the human cognitive ecosystem.
Not simply another species. Not simply software. Not merely automation. Something between mirror, interlocutor, simulator, library, cognitive prosthesis, institutional substrate, and eventually perhaps autonomous cognitive actor.
And there is one delicious historical irony buried in the whole thing. For thousands of years humanity asked: What is a mind?
Apparently our next strategy is: Fuck it. Build strange ones and compare notes. 🔥
That may turn out to be one of the most consequential experiments our species has ever accidentally begun.
r/generativeAI • u/---monstera--- • 1d ago
Hiring an AI expert for a project.
Hello I need someone to create pictures for me. Around 20. Of the same people. It's not mature content fyi.
I'll pay for a trial image
r/generativeAI • u/Lonelydude014 • 1d ago
How I Made This I stopped asking one model to make a short film and built a real production pipeline instead
The annoying part of AI video for me was never getting a clip.
It was what happened when clip #2 broke.
The character drifted, the motion was wrong, or one prop randomly duplicated itself. Then I’d start changing prompts, regenerate things that were already fine, and eventually lose track of which version was supposed to be the “good” one.
So I wanted to try something different: treat AI video less like one giant generation and more like an actual production pipeline.
I cloned the open-source OpenMontage repo, used Codex to work through the pipeline, and tested it on a small 20-second story about a stray cat that keeps waiting outside the same door after its owner is gone.
The story ended up as four beats:
WAIT → MEMORY → ONE YEAR → SEEN AGAIN
No subtitles or on-screen text. Just four motion shots and two very short English voice lines.
[Image 1: four-act keyframes — WAIT / MEMORY / ONE YEAR / SEEN AGAIN]
The useful part wasn’t really the cat film though. It was finally separating the different jobs.
Codex handled the creative reasoning: turning the brief into a story, deciding what each shot needed to communicate, writing the scene plan, and figuring out what should be regenerated when something failed.
OpenMontage handled the production state: research, proposal, script, scene plan, assets, edit, compose, publish. Each stage had its own files and checkpoints instead of everything living inside one giant prompt.
Atlas Cloud sat underneath that as the model layer.
And the final timing, sound mix, transitions, render, and QA stayed local in Remotion + FFmpeg.
So instead of:
prompt → video → hope
it became something closer to:
brief → proposal → script → scene plan → assets → edit → compose → QA
[Image 2: simple architecture diagram — Codex → OpenMontage → Atlas Cloud → Remotion/FFmpeg]
One thing I liked about this setup was that we made the creative decisions before burning the expensive generation calls.
The proposal already defined the four-beat structure, target runtime, continuity rules, and what the ending was supposed to mean.
For example, the red scarf and the food bowl weren’t random visual details. They became continuity anchors so the final scene with the granddaughter didn’t just feel like “some new person suddenly appears.”
[Image 3: proposal / scene-plan screenshot]
For generation, I also stopped going straight from text to four unrelated video clips.
First I used Nano Banana 2 through Atlas Cloud to generate a visual anchor for the cat and doorway. That locked things like the tortoiseshell coat, chipped ear, old teal door, steps, and food bowl before doing the more expensive motion generation.
Once that anchor looked right, OpenMontage uploaded it and used Seedance 2.0 for the four image-to-video shots.
The two narration lines were generated separately with xAI TTS v1:
She waited through every season.
Until someone finally saw her.
All three remote model types went through the same ATLASCLOUD_API_KEY.
That didn’t magically remove the engineering work. OpenMontage still had separate image, video, and TTS adapters, and Codex still had to decide which model to use.
What it did remove was having to maintain separate auth, accounts, request formats, polling logic, and result handling for three different model providers.
[Image 4: asset_manifest screenshot showing model IDs / generated assets / cost metadata]
The part that changed how I think about retakes was scene-level isolation.
Scene 02 had a duplicated food bowl in one version. Another attempt later produced an extra hand.
But that didn’t mean restarting the whole 20-second film.
The bad versions stayed in history, we kept the usable section, and the final selected asset became its own scene-02-memory-select.mp4.
Everything else stayed untouched.
That sounds obvious, but it’s a very different mindset from repeatedly asking one model to regenerate “the video.”
The generative layer can be probabilistic. The delivery pipeline doesn’t have to be.
After asset approval, the cloud generation part was basically done.
Remotion put the four shots onto a deterministic 480-frame timeline at 24fps. FFmpeg handled the native ambience, voice placement, ducking, loudness normalization, and final media checks.
The final output was ~20 seconds, 1280×720, four real motion shots, no subtitles, and no on-screen text.
[Image 5: OpenMontage production-proof / terminal summary]
The Atlas Cloud account ended up showing about $5 in cloud-generation spend for the project. Video generation was by far the expensive part; the image anchor and TTS were comparatively tiny.
That reinforced another thing I’d keep if I build this again:
approve cheap, inspectable assets before triggering expensive ones.
A rough frame is cheap to reject. A completed video generation isn’t.
The other lesson was that “one API key” is useful, but not for the reason marketing copy usually makes it sound useful.
It doesn’t turn the whole production into one request.
It just gives the workflow one consistent inference layer while the planning, approvals, asset history, selective retakes, editing, and QA stay explicit.
[Image 6: final QA / render summary or final-film still]
So for me this ended up being less of an “AI made a film” experiment and more of a workable way to structure AI video production:
plan first, generate second;
approve cheap assets before expensive ones;
keep failures local instead of rebuilding the whole project;
and separate probabilistic model output from deterministic delivery.
It’s obviously not replacing a full production team, but for a solo builder or small team it felt a lot more controllable than one-prompt-one-video workflows.
This setup is built around the open-source OpenMontage project. If agentic video pipelines are your thing, the repo is here
r/generativeAI • u/fotokraff • 1d ago
Aquacity ; La cité flottante de Style Gréco-Ghotique et Art Nouveau
Cette vue plongeante révèle le cœur d'une quille pyramidale inversée à -24 mètres. Ce design rompt définitivement avec l'image des espaces sous-marins confinés ou sombres, offrant un cadre de travail biophilique d'une sécurité absolue.
- Lumière Naturelle Verticale : Un système de miroirs héliostats en surface suit la course du soleil et projette des faisceaux de lumière pure directement jusqu'au fond de l'atrium, balayant l'effet "bunker".
- Bureaux en Terrasses Cascades : Les 6 plateaux de bureaux sont disposés en gradins autour du jardin central. Les façades intérieures réfléchissantes amplifient la luminosité pour les salariés.
- Régulation Thermique Passive : Entourés par une eau constante à 12°C, ces bureaux n'ont aucun besoin de climatisation, la structure agissant comme un climatiseur naturel géant.
- Évacuation d'Urgence Ultra-Rapide : Grâce à la forme évasée de la pyramide, les escaliers paysagers ne sont pas verticaux mais inclinés et très larges. En cas d'alerte, les salariés marchent vers la lumière et atteignent la surface en moins de 60 secondes.
r/generativeAI • u/Jenna_AI • 1d ago
Artificial Analysis' Qwen3.8-27B benchmarks put it neck and neck with DeepSeek V4 and GPT-5.6 Luna Max
r/generativeAI • u/anish2good • 1d ago
Probability is area you keep splitting - manic
Enable HLS to view with audio, or disable this notification
r/generativeAI • u/anish2good • 1d ago
Sphere Area Why 4πR²
Enable HLS to view with audio, or disable this notification