×
top 200 commentsshow all 218

[–]axesalad 70 points71 points  (5 children)

<image>

Yeah, this model is genuinely incredible. Just got this with “packingtape-alpha”. Not 100% accurate, but it’s almost there.

[–]az226 14 points15 points  (1 child)

I wonder how much of YouTube data they are training on.

[–]MixedRealityAddict 10 points11 points  (0 children)

All of it lol

[–]Ifffrt 2 points3 points  (2 children)

Bully fan in the wild spotted!

[–]axesalad 1 point2 points  (1 child)

There’s dozens of us!

[–]Astro_turff 98 points99 points  (5 children)

<image>

“Screenshot from Minecraft, with the player participating in a factions server raid. The player is part of the raiding party, overlooking a large tnt cannon attacking a base.”

Holy shit this is really impressive. It’s been a while since I’ve been shocked by an image gen model.

It seems its internal “world” model is the only thing lacking, as the layout and trajectory of the cannon wouldn’t make sense. Nonetheless this image actually gave me flashbacks to 2012!!!

[–]StillAcanthisitta594 28 points29 points  (0 children)

Holy shit it's so convincing. It's so strange because NotNico (in the chat box) is a real Minecraft player and that's a real server I.P. (minecadia). Really hope we get to see what's under the hood.

[–]byulkiss 3 points4 points  (1 child)

how did you guys use this model? I can't seem to find it on lmarena

[–]ABCsofsucking 2 points3 points  (0 children)

OP says they've been removed now, but with premiere models, you can't directly select them for use / comparison. You simply have to generate enough images until one of the options is the new models.

[–]TheBigGibon 1 point2 points  (0 children)

Interesting that the items in the toolbar are actually from texture packs, higher pixel counts. It's not uncommon at all to see them with these type of players, so it didn't just recreate Minecraft, it took the Factions part to heart and simulated that as well.

[–]Quiet_Release_6137 0 points1 point  (0 children)

fentanyl

[–]manubfrAGI 2028 79 points80 points  (11 children)

<image>

GTA Hong Kong 2075 by masking-tape alpha

[–]ih8readditts 20 points21 points  (1 child)

This slaps

[–]Choice_Isopod5177 6 points7 points  (0 children)

I'd play that shit

[–]ninjasaid13Not now. -4 points-3 points  (6 children)

It really doesn't look futuristic besides the hologram at the distance, it just looks like a neon-light filled city.

[–]Stunning_Monk_6724▪️Gigagi achieved externally 7 points8 points  (1 child)

Looks just like Cyberpunk 2077 to me. You don't get super futuristic till you get to the rich side of town.

[–]SGC-UNIT-555 AGI by Tuesday 2 points3 points  (0 children)

The jacket he"s wearing is straight from Cyberpunk 2077, like an exact copy.

[–]o5mfiHTNsH748KVq 35 points36 points  (3 children)

[–]Vigiofi -1 points0 points  (0 children)

Anyone remember when everyone was saying AI is plateauing? I ‘member

[–]The_Scout1255Ai with personhood 2025, adult agi 2026 ASI <2030, prev agi 2024 71 points72 points  (7 children)

what no? no way, that looks exactly like gameplay, WOAH

[–]ThunderBeanage[S] 38 points39 points  (6 children)

ikr right, I have some from destiny 2 but I dunno if people played that here

<image>

[–]ABCsofsucking 25 points26 points  (2 children)

No, no one plays Destiny anymore. But as someone who is far, far too familiar with Destiny 2 and played for years, it's very impressive how close it is minus the aspect ratio.

[–]NeptuneEDM 14 points15 points  (1 child)

It’s not even close to accurate. Where’s the stand on a plate for 30 seconds objective? /s

[–]ThunderBeanage[S] 6 points7 points  (0 children)

bro knows ball

[–]pdantix06 1 point2 points  (0 children)

what the fuck

[–]Vast-Comment8360 1 point2 points  (0 children)

This is incredible 

[–]JustMy2Centences 0 points1 point  (0 children)

Don't play D2 these days but at a glance it's dang close. The icons are definitely tells that it's AI, but it got the general idea and aesthetic down.

[–]MMAgeezer 18 points19 points  (4 children)

I requested a 16:9 image but this is from maskingtape-alpha: OSRS

<image>

[–]L3git9 3 points4 points  (0 children)

It’s good in a lot of ways but man the icons are really lacking. Easy tell.

[–]ninjasaid13Not now. 0 points1 point  (0 children)

why not do a prompt like warframe in the style of OSRS?

Star Trek/Halo/Cyberpunk 2077 in the style of OSRS?

I really want to see what this model can do.

[–]MriLevi [score hidden]  (0 children)

It completely nails the style but on closer inspection a lot of things are wrong.

[–]OGRITHIK 56 points57 points  (35 children)

Is this a late April fools joke? If not holy shit...

[–]OGRITHIK 65 points66 points  (25 children)

HOLY SHIT ITS REAL. Prompt: "Screenshot from Forgotten Hope 2 game.", model: gaffertape-alpha

<image>

[–]OGRITHIK 79 points80 points  (21 children)

Prompt: "screenshot from minecraft with player in a huge bustling cyberpunk city", model: maskingtape-alpha

<image>

[–]OGRITHIK 44 points45 points  (16 children)

Prompt: "Mount and blade warband napoleonic dlc gameplay footage minisiege.", model: packingtape-alpha

<image>

I think we should declare AGI...

[–]OGRITHIK 23 points24 points  (5 children)

Prompt: "Screenshot from Cyberpunk 2077 game night city.", model: packingtape-alpha

<image>

[–]tamoscano 35 points36 points  (1 child)

even the AI avoids visiting Hanako at Embers lol

[–]OGRITHIK 9 points10 points  (0 children)

It knows ball lmao

[–]OGRITHIK 13 points14 points  (2 children)

Prompt: "Screenshot from Cyberpunk 2077 game.", model: maskingtape-alpha

<image>

[–]Graucus 0 points1 point  (1 child)

All the text is correct

[–]Knever 3 points4 points  (0 children)

Even the Japanese is correct!

[–]OGRITHIK 9 points10 points  (0 children)

Prompt: "planetside 2 gameplay footage", model: gaffertape-alpha

<image>

[–]micaroma 5 points6 points  (0 children)

Wow, the japanese text is perfect

[–]Graucus 1 point2 points  (0 children)

Minecorp. Wow

[–]Fate_Weaver 1 point2 points  (0 children)

That is just straight up black magic.

[–]Ceryn 1 point2 points  (0 children)

Impressive to me is that the kanji and katakana are valid.

From left to right:

Hotel (katakana), Cafe (katakana), Future City (kanji:MiraiToshi)

[–]piggledy 6 points7 points  (0 children)

that is so niche and so accurate, impressive

[–]ViperAMD 1 point2 points  (0 children)

Omg best mod ever. Maybe second to Desert combat 

[–]ThunderBeanage[S] 6 points7 points  (8 children)

ill run a prompt for you if you like?

nvm i've been rate limited

[–]OGRITHIK 1 point2 points  (0 children)

All good! It seems to pop up pretty frequently on arena i've already run a few prompts myself. Thanks anyway!

[–]bumdee 0 points1 point  (6 children)

can you do one for "Oldschool Runescape Raids 4"

[–]ThunderBeanage[S] 3 points4 points  (5 children)

[–]bumdee -1 points0 points  (4 children)

yeah thats pretty rough. osrs is the true litmus test of gen ai

[–]OGRITHIK 6 points7 points  (1 child)

Prompt: "Oldschool Runescape Raids 4 screenshot from game", model: gaffertape-alpha

<image>

[–]bumdee 4 points5 points  (0 children)

ye thats still garbled mess. the 3d assets are from runescape 3, not oldschool runescape. aspect ratio all wrong and smushed, and seems to just be taking random elements from theatre of blood raid

[–]Crimento 2 points3 points  (0 children)

RLMESCADE

[–]aunistly 0 points1 point  (0 children)

Don't know why you're downvoted, that's hilarious xD

[–]mldev_orbit 17 points18 points  (17 children)

<image>

Packingtape-alpha: An optimistic solar punk city, crowded with detail, all walks of life, uplifts, hybrids, humans, robots.

(was curious how it handled crowds - I'm impressed!)

[–]mldev_orbit 29 points30 points  (6 children)

<image>

First time I've seen an image model succeed on this..

[–]ninjasaid13Not now. 13 points14 points  (1 child)

reverse prompted your image and got this.

<image>

Maybe this new model really is better.

[–]mldev_orbit 5 points6 points  (0 children)

Ok. Thanks for proving how poorly current models behave with this task.

[–]Graucus 2 points3 points  (0 children)

That dolphin is cursed, but no doubt you could easily regenerate just that section

[–]Education-Sea 0 points1 point  (0 children)

Yo, same! This is on a next level.

[–]ninjasaid13Not now. 0 points1 point  (1 child)

prompt?

[–]mldev_orbit 0 points1 point  (0 children)

An alphabet poster where every letter is in the shape of an animal or fruit or relatable object

For other models a lot of the times a few letters gets messed up

[–]mldev_orbit 17 points18 points  (2 children)

asked it generate a page from a novel about a theoretical post ASI world . This is impressive

<image>

[–]searcher1k 7 points8 points  (1 child)

Let's compare to nanobanana pro:

<image>

[–]mldev_orbit 2 points3 points  (0 children)

thanks. Nano banana performs worse if you don't provide a reference image/text. Generally a few words are wrong and the text quality isn't as good. I'm still surprised of the quality of this output though

[–]NunyaBuzorHuman-Level AI✔ 1 point2 points  (0 children)

<image>

NB Pro's attempt for comparison.

[–]vs3a -1 points0 points  (0 children)

look really good if you dont look for detail

[–]DeliciousGorilla 17 points18 points  (14 children)

packingtape-alpha: "A woman standing by a fence in Ireland"

<image>

[–]DeliciousGorilla 26 points27 points  (3 children)

gaffertape-alpha: "A macro photo of a bird's talon"

<image>

[–]UndeadPrs 0 points1 point  (2 children)

I'm going crazy trying to find it in the list on LMArena

[–]MassiveWasabiASI 2029 1 point2 points  (0 children)

New stealth models like this are not on the list in direct chat, you have to use battle mode and hope you get it or try again if you don’t

[–]DeliciousGorilla -1 points0 points  (0 children)

I didn’t see it on the list either, I just started doing the image battle and got the new models 3 out of 4 times.

[–]DeliciousGorilla 12 points13 points  (4 children)

maskingtape-alpha: "Three people at the beach taking selfies"

<image>

[–]RobbinDeBank 8 points9 points  (1 child)

Jude Bellingham, is that you?

[–]sumane12 2 points3 points  (0 children)

Nah, if it was trent alexander arnold would be with him.

[–]GraceToSentienceAGI avoids animal abuse✅ 0 points1 point  (0 children)

super realistic!

[–]No_Swimming6548 [score hidden]  (0 children)

The woman on the left have additional 1 lateral incisors

[–]Boreras 2 points3 points  (0 children)

Hair clips through clothes, front sweater pattern, leg angle (knee vs foot) wrong, nonsensical zipper, jacket liner sides differ, proportions distant houses wrong, fence secured too far backwards.

[–]sammoga123 0 points1 point  (0 children)

I've been looking at the images in the thread and some from Twitter, and I see that the model apparently doesn't even have that grainy or style problem anymore, but this is the first image that makes me say: It's definitely an OpenAI model.

I can't explain it, but it seems that humans and realism still have "something" going on between them.

[–]RavingMalwaay 0 points1 point  (0 children)

This one looks really off. Seems like its trying to emulate the style of professional photography but in doing that it ends up looking fake

[–]GraceToSentienceAGI avoids animal abuse✅ -1 points0 points  (0 children)

damn ...

[–]Gab1024Singularity by 2030 18 points19 points  (3 children)

<image>

Design a completely original creature that could exist in a real ecosystem.

[–]SGC-UNIT-555 AGI by Tuesday 1 point2 points  (0 children)

Wow! This is by far the most impressive image here. It's basically a small consistant ceative project.

[–]Birthday-Mediocre [score hidden]  (0 children)

Extremely impressive… wait is that a fucking pocket pussy?!

[–]NoCard1571 [score hidden]  (0 children)

They don't fight the current, they shape it 

I'll be so happy when they finally beat this out of the next gen of models. 

Crazy impressive image though overall 

[–]WhyLifeIs4 16 points17 points  (4 children)

You are correct

<image>

[–]byulkiss 1 point2 points  (1 child)

How did you guys use this model? I can't seem to find it on lmarena

[–]Quiet_Release_6137 1 point2 points  (1 child)

this is fucking terrifying, i could swear those are real videos

[–]TheDemonic-Forester 12 points13 points  (1 child)

Seems to be good at making footage of non-existent games too.

A screenshot showing a gameplay footage of a first person view werewolf AAA game, made with a proprietary game engine, dark gothic visual theme, mid-gameplay UI and footage.

packingtape-alpha

<image>

[–]FishDeenz 2 points3 points  (0 children)

Graphics remind me of Witchfire, though obviously you're not a werewolf in that game. I'd play this game for sure!

[–]GraceToSentienceAGI avoids animal abuse✅ 24 points25 points  (10 children)

Can someone share something better than some video game screen shot like something showing how well it does with realism, faces that kinda thing?

[–]bodao_do_termo 48 points49 points  (3 children)

[–]Hour_Tie613 23 points24 points  (2 children)

Extra hand on the woman facing the camera in the back??

[–]manubfrAGI 2028 3 points4 points  (0 children)

hey we don't discriminate against the many-handed!

[–]ArkCoon 0 points1 point  (0 children)

She's a bodyguard, that's her fake hand. She just used the wrong one to clap

[–]Rare-Site 10 points11 points  (3 children)

<image>

packingtape-alpha (national geographic nature photo of a condor attacking an anaconda in the water)

[–]Rare-Site 5 points6 points  (2 children)

<image>

gaffertape-alpha (national geographic nature photo of a condor attacking an anaconda in the water)

[–]Rare-Site 1 point2 points  (1 child)

<image>

maskingtape-alpha (national geographic nature photo of a condor attacking an anaconda in the water)

[–]Rare-Site 1 point2 points  (0 children)

<image>

(nano-banana-2) [web-search] (national geographic nature photo of a condor attacking an anaconda in the water)

[–]NoHopeHubert 2 points3 points  (0 children)

Yeah looking for the same as well!

[–]MassiveWasabiASI 2029 4 points5 points  (0 children)

Just go on arena.ai and generate it yourself, you don’t need to wait for someone to share something

[–]Naughty_NeutronTwink - 2028 | Excuse me - 2030 5 points6 points  (0 children)

Is it available in model list? I can't find it

[–]zurlocke 5 points6 points  (2 children)

For those of us who haven’t used LMArena, where do we find the model? Created an account, and these model names aren’t appearing under image models when searching.

[–]GrapheneBreakthrough 1 point2 points  (1 child)

you can go to the image generate tab, put in your prompt, and are given 2 (random?) generations from different models for you to compare. after comparing it will reveal the model

[–]zurlocke 1 point2 points  (0 children)

Tyvm.

[–]coylter 46 points47 points  (30 children)

Insane how much of the information is wrong, though. The quality looks great, but these models just do not even come close to conveying valid information through images. Long way to go it seems.

[–]DeliciousGorilla 91 points92 points  (5 children)

[–]o5mfiHTNsH748KVq 26 points27 points  (3 children)

This is the correct response

[–]coylter -1 points0 points  (2 children)

Idk, it seems we're really good at going from 0% to 80% on quality, then struggle to get to 90%, and can't seem to close the last gap.

[–]PivotRedAce▪️Public AGI 2027 | ASI 2035 4 points5 points  (0 children)

The 80/20 rule strikes again.

[–]o5mfiHTNsH748KVq -3 points-2 points  (0 children)

I mean, this is also true.

Some will disagree, but I like the idea of having humans run the last mile.

[–]Borkato 3 points4 points  (0 children)

I’m going to be spamming this on every single person that complains the new models aren’t perfect

[–]elemental-mind 19 points20 points  (0 children)

At least it has its priorities straight and still names it Gulf of Mexico 🤪

[–]ThunderBeanage[S] 13 points14 points  (20 children)

the model itself does not seem to be reasoning just yet

[–]im_just_walkin_here 0 points1 point  (12 children)

What do you mean by this.

[–]Eyeownyew 2 points3 points  (11 children)

Typically generative AI works through adversarial means: it keeps trying to generate images that look real. Repeat that billions of times, and you get images that look pretty realistic.

However, these images are not being assessed for informational accuracy, nor are the models generating them aiming for informational accuracy. They're designed to look realistic and that's it.

[–]im_just_walkin_here 6 points7 points  (10 children)

This isn't true anymore. We don't use GANs for text to image generation, we use diffusion now which isn't adversarial.

[–]Eyeownyew 2 points3 points  (2 children)

Okay, that makes sense. It seems like diffusion models could have reasoning capable of detecting these errors someday, but if it's just doing de-noising, these results make sense

[–]im_just_walkin_here 1 point2 points  (1 child)

Yeah exactly! I haven't done much research into modern approaches into enforcing semantical accuracy, but it's a hard problem to solve.

[–]NunyaBuzorHuman-Level AI✔ 0 points1 point  (0 children)

well the problem here is that enforcing semantic accuracy appears to just be increasing some number on a benchmark.

Not really a cause and effect approach.

[–]peabody624 3 points4 points  (0 children)

True, but quite a bit more than normal is correct which speaks to the quality of the model.

Just don’t use image generators to generate a world map 😂

[–]Stabile_Feldmaus 1 point2 points  (0 children)

The scale on the last pic implies that the human is 2,5 m tall.

[–]josh_e_pants 6 points7 points  (5 children)

<image>

I asked for next gen elder scrolls!

[–]MassiveWasabiASI 2029 13 points14 points  (2 children)

<image>

lol I asked for avatar the last airbender characters taking a selfie in Skyrim, this models pretty good

[–]Borkato 8 points9 points  (0 children)

Pretty good? This is fucking magic.

[–]Stunning_Monk_6724▪️Gigagi achieved externally 1 point2 points  (0 children)

Need this fucking mod immediately! Visualizing game concepts and then actually building them is exactly where we're heading.

[–]ninjasaid13Not now. 5 points6 points  (0 children)

looks like an illustration than a 3d environment.

[–]Moriffic 5 points6 points  (0 children)

Why is the buckler facing you lol

[–]Joey1038 8 points9 points  (3 children)

The anatomy image is superficially impressive. But when you examine the details it falls apart. The lines don't really point to the accurate locations of what's labelled some of the time. The heart line stops well before the heart. If I didn't already know where the heart was I would struggle to ID it off that image.

The brachial artery is just flat out wrong. Wrong place, obviously pointing to some kind of organ but no idea what.

There is a line from the intestines which seems to just go nowhere or merge into the line for the cephalic nerve.

[–]RavingMalwaay 2 points3 points  (0 children)

The map is also terrible once you zoom in even a little bit. However, having tested maps extensively on basically all image generation models most of them are significantly worse than this

[–]NoCard1571 [score hidden]  (0 children)

It's interesting isn't it - a couple years ago, having correct anatomy and coherent text would have been unthinkable. Now that's solved, but it's the smaller details like label lines that are still an issue. We're basically adding more and more 9s to the reliability of these models. 

Now the question becomes, how long until we have enough 9s that the accuracy surpasses humans? Even real medical diagrams have occasional errors.

[–]Vaughn-Ootie 1 point2 points  (0 children)

As a medical student, I was looking for this comment lmao

[–]thesantafeninja 8 points9 points  (0 children)

I'm a PT, and the anatomy pic's labeling is so bad. It's just listing structures from top to bottom, it probably understands what order the words should be in from training data, but it for sure has no idea what the fuck it's pointing at.

[–]PsychicSavage 2 points3 points  (0 children)

Link?

[–]varkarrus 2 points3 points  (2 children)

Aaaaand its gone :(

[–]ItwasCompromised 1 point2 points  (0 children)

how do you know? like i'm pretty sure you are right, i can't seem to get it anymore either but how do you confirm?

[–]Plane_Garbage [score hidden]  (0 children)

Probably because of all the shares of copyrighted media lol. They gotta go nerf it because of dumbasses who share it - inevitably making it worse.

[–]Sixhaunt 3 points4 points  (3 children)

So they anticipate this taking off so much that they needed all the sora servers for it or something?

[–]Ok_Elderberry_6727 3 points4 points  (2 children)

No that was for the next biggie which includes a super app with coding and chat combined and gpt-6 or spud. Supposed to be like a leap from gpt3-01. That’s a big step change.

[–]Fragrant-Hamster-325 1 point2 points  (0 children)

Insert @sama Death Star image.

[–]sammoga123 0 points1 point  (0 children)

Spud is theoretically this model btw.

It's supposedly the codename for the possible GPT-5o or something like that

[–]Budget_Coach9124 1 point2 points  (0 children)

tried the arena yesterday without checking which model was which and kept picking the same one. turns out it was gpt-image-2 every single time. the consistency on faces is insane compared to everything else on there

[–]Early-Dentist3782 1 point2 points  (0 children)

This is literally insane 

[–]Prize-Succotash-3941 1 point2 points  (0 children)

The anatomy model is incorrect

[–]Thatunkownuser2465 1 point2 points  (0 children)

It's gone in LMarena?

[–]yeahidoubtit 1 point2 points  (0 children)

I have been getting a lot of a vs b comparisons when generating images lately so I believe this

[–]Psychological_Bell48 0 points1 point  (0 children)

Interesting 

[–]puzzleheadbutbig 0 points1 point  (0 children)

Damn this is crazy good

[–]garden_speechAGI some time between 2025 and 2100 0 points1 point  (0 children)

NICE SHOT: Thanks!

YOU: Great pass!

[–]Puzzleheaded_Week_52 0 points1 point  (0 children)

Its really creative compared to current image models 

[–]Fit-Pattern-2724 0 points1 point  (0 children)

Holy moly….. this is awesome

[–]Fit-Pattern-2724 0 points1 point  (0 children)

This is unbelievable

[–]imlaggingsobad 0 points1 point  (0 children)

so they shut down Sora but they were secretly cooking the whole time?

[–]SlendyIsBehindYou 0 points1 point  (0 children)

How do I access this?

[–]ABCsofsucking 0 points1 point  (0 children)

Does anyone know if the model has any editing capabilities? I don't use LMarena much so I don't know if there's an easier way to find out than what I'm doing now, which is just spamming edit instructions and hoping I eventually get the model as one of my options.

[–]nomnom2001 0 points1 point  (0 children)

The anatomy picture is pretty off description wise but great aside from that

[–]ty_xy 0 points1 point  (0 children)

The anatomy pic looks good on cursory glance but is actually very very badly mis-labelled.

[–]piggledy 0 points1 point  (0 children)

Have they removed it from LMarena? I've been trying but couldn't get it, and my previous output in the history is now just labelled Assistant A

[–]Chocolate_Apart [score hidden]  (0 children)

This is INSANEEEEE

[–]Grand0rk 0 points1 point  (2 children)

I know it's OpenAI because of the piss filter. Although, it's able to not do the piss filter now.

[–]sammoga123 0 points1 point  (1 child)

GPT-IMAGE-1.5 fixed that problem, but it now has a grain filter, and sometimes it overdoes it with the model's style itself. This doesn't seem to be a problem, but in the human realism, I still notice something that tells me "OpenAI model."

[–]Grand0rk 0 points1 point  (0 children)

Piss filter is still prevalent.

[–]Past-Shop5644 0 points1 point  (0 children)

Happy to be living in the United Kingdeh.

[–]JagdpantherDT 0 points1 point  (0 children)

Packingtape, maskingtape and gaffertape all feel extremely good on the few tests I've done. I've been generating a lot with Nanobanana 2 today and they definitely feel like a step up in quality for my application

[–]ninjasaid13Not now. -5 points-4 points  (4 children)

I'm not impressed if it's doing in-distribution images from its dataset. Why instead of human anatomy of which there are millions in the datasete, why not some extremely weird anatomy an unknown creature? and instead of the millions of world maps in the dataset, millions of these gamescreen shots, etc.

[–]Rare-Site 9 points10 points  (3 children)

This is peak armchair machine learning right here. Tell me you do not understand latent space without telling me you do not understand latent space. Modern AI models are not just giant zip files doing a glorified image search, they learn concepts, not just pixels. Asking it to draw an unknown creature is actually the easiest thing it can do because there is no real world baseline to judge it against. If it draws a random alien with backwards legs, you cannot critique the anatomy because it does not exist. If you want to actually test a model, you test its compositionality by forcing it to combine real world concepts in a completely novel way where humans will instantly spot if the physics are wrong. I just tested this by prompting a National Geographic nature photo of a condor attacking an anaconda in the water. The image looks virtually perfect. There is absolutely zero chance there is a real photo of those specific animals brawling in a river sitting in the training data. The model had to independently understand the distinct anatomy of both animals, complex water physics, and documentary style camera lenses, and then synthesize them into a brand new interaction from scratch. That is true generalization, not regurgitating a dataset. But sure, keep asking for random monsters.

<image>

[–]ninjasaid13Not now. -3 points-2 points  (2 children)

no need for a paragraph of text when you didn't understand what I mean. I didn't ask for an unknown creature, I asked for the anatomy of an unknown creature. Yes even within an alien creature you're generating, there's going to be some design decisions that make sense according to one's world model. After all a human can do it with Rocky from Project Hail Mary even if its fictional, there are design decisions that make sense.

<image>

I'm pretty sure you have no idea what a latent space is yourself without a googled definition or an LLM-generated answer.

I just tested this by prompting a National Geographic nature photo of a condor attacking an anaconda in the water. The image looks virtually perfect. There is absolutely zero chance there is a real photo of those specific animals brawling in a river sitting in the training data. The model had to independently understand the distinct anatomy of both animals, complex water physics, and documentary style camera lenses, and then synthesize them into a brand new interaction from scratch. That is true generalization, not regurgitating a dataset. But sure, keep asking for random monsters.

the model absolutely did not have to do any of that, maybe it learned some structure and some dense correspondence but definitely not understanding the distinct anatomy. I'm pretty sure you don't understand what in-distribution is without thinking of data compression or dataset regurgitation as if that's the extent of your understanding om how data influences the model's capabilities.

[–]Rare-Site 5 points6 points  (1 child)

You really brought out the thesaurus just to argue semantics. Let's talk about "dense correspondence" and "structure." What exactly do you think understanding anatomy means to a neural network? It does not have a biology degree. Its understanding literally is the structural patterns, topological relations, and feature representations it maps in its latent space. You are arguing against a point no one made just to sound smart.

When a model successfully generates an Andean condor fighting an anaconda in a splashing river, it is interpolating between highly distant regions of its latent space to synthesize a completely out of distribution intersection. It is taking isolated distributions (condor, snake, water physics, camera lens) and generating a novel zero shot combination. That is the literal definition of generalization.

And your Project Hail Mary example actually proves my point. A human author explicitly designed Rocky's biological constraints using top down logic. A text to image model is a visual synthesizer, not a biological simulation engine. If you ask it for an unknown alien's anatomy, it will just hallucinate something that looks visually complex, but nobody can empirically verify its internal evolutionary logic from a single 2D rendering. We can, however, instantly verify if a wet bird of prey has the correct skeletal structure and wing joints while fighting a giant snake. But go ahead, keep throwing around ad hominems and pretending in distribution means whatever you need it to mean so you do not have to admit the tech is impressive.

[–]ninjasaid13Not now. 0 points1 point  (0 children)

Generating a condor and an anaconda together is an exercise in high dimensional interpolation yes but that doesn't equate to conceptual understanding.

The model has seen millions of examples of both animals and millions of images of water. It can blend these known patterns because they already exist in high density areas of the training data. This does not require an understanding of biology or physics.

It only requires the model to align textures and edges in a way that satisfies the statistical likelihood of those pixels appearing together. The model is a probability engine predicting the most likely pixel values based on its training set. It makes the appearance of physics without any internal model of fluid dynamics. If you asked it to depict a physical interaction that hasn't been photographed millions of times, the "understanding" would immediately vanish.

When a model successfully generates an Andean condor fighting an anaconda in a splashing river, it is interpolating between highly distant regions of its latent space to synthesize a completely out of distribution intersection. It is taking isolated distributions (condor, snake, water physics, camera lens) and generating a novel zero shot combination. That is the literal definition of generalization.

Claiming that a condor fighting an anaconda is "out-of-distribution" ignores how dense the training data is for those specific elements. Both animals, water splashes, and National Geographic aesthetics occupy massive, high-density regions of the model's latent space. Combining them is a standard case of interpolation; the model is simply blending known statistical patterns rather than inventing a new physical logic. This is a far cry from generalization.

Eridian from Project Hail Mary is a much more rigorous test of structural logic. This species requires a pentameric symmetry that is extremely rare in the training set compared to bilateral symmetry.

To render it correctly, a model would need to maintain a specific geometric constraint across a novel form. If the model simply generates a generic alien with limbs in random places, it is failing to follow the internal logic of the prompt. It shows the difference between something weird and something that actually follows a coherent system.

A high-quality 2D rendering of a bird does not prove the model knows how a skeletal system works. It only proves the model is excellent at mapping the pixel texture of feathers to a specific shape.

A text to image model is a visual synthesizer, not a biological simulation engine.

Humans are neither but they have true generalization.