Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Stable Diffusion is mind-blowingly good at some things. If you are looking for modern artistic illustrations (like the stuff that you would find on the front page of Artstation) - it's state of the art, better in my opinion then Dalle-2 and Midjourney.

But, the interesting thing is that while it is so good in producing detailed artworks and matching the styles of popular artists, it's surprisingly weak at other things, like interpreting complex original prompts. We've all seen the meme pictures made in Craiyon (previously Dalle-mini) of photoshop-collage-like visual jokes. Stable Diffusion with all its sophistication is much worse at those and is struggling to interpret a lot of prompts that the free and public Craiyon is great with. The compositions are worse, it misses a lot of requested objects or even misses the idea entirely.

Also as good as it is at complex artistic illustrations, it is as bad at minimalistic and simple ones, like logos and icons. I am a logo designer and I am already using AI a lot to produce sketches and ideas for commercial logos, and right now the free and publicly available Craiyon is head and shoulders better at that then Stable Diffusion.

Maybe in the future we will have a universal winner AI that is the best at any style of pictures that you can imagine. But right now we have an interesting competition when different AI have surprising strengths and weaknesses and there's a lot of reason in trying them all.



Just think where we'll be two more papers down the line


For those unaware this is a catchphrase of Dr. Karoly Feher from the absolutely wonderful YouTube channel "Two Minute Papers" which focuses on advances in computer graphics and AI.


Random rant: it feels like over time Two Minutes Paper has started to lean more and more into its catchphrases and gimmicks, while the density of interesting content keeps decreasing.

The whole "we're all fellow scholars here" bit feels like I'm watching a kid's show about science vulgarization, patting me on the head for being here.

"Look how smart you are, we're doing science!"

I dunno. I like the channel for what it is (a vulgarization newsletter for cool ML developments) but sometimes the author feels really patronizing / full of himself.


I agree that I like to for what it is - something more along the lines of Popular Science or Wired than Scientific American if you want to compare to magazines. However, the content, while surface level, is always accurate - something that can’t be said for other content creators in the field.


I agree that it can be a lot at times, especially if you watch several in a row, but I dunno, I kind of love that he's keeping that enthusiasm (real or not). I think the world is a brighter place because of it. Just a tiny bit, but still.


I think the biggest benefit is the curation aspect. After all, how much can you actually learn in two minutes? Once I see something interesting, I go and read through the actual paper. Having said that, you're lucky if you can find a paper with enough details to actually reproduce the work.



You're mistaking earnest for patronizing. He's a genuinely positive dude.


he stopped summarizing methods at some point- now its just results


> Now squeeeze those papers!


it's surprisingly weak at interpreting complex original prompts because the model is really small, the text encoder is just 183M parameters. Craiyon is much larger.


I have a penchant for wanting to make technically "bad" or heavily stylized photos - and Stable Diffusion is pretty poor at those. There's very little good bokeh or tilt shift stuff and CCTV/Trailcam doesn't come out too well.

In fact Dall-E isn't as impressive for some styles as "older" models (Jax/Latent Diffusion etc)


My hunch is that is the result of this: https://github.com/CompVis/stable-diffusion#weights

> 515k steps at resolution 512x512 on "laion-improved-aesthetics" (a subset of laion2B-en, filtered to images with an original size >= 512x512, estimated aesthetics score > 5.0

https://github.com/LAION-AI/laion-datasets/blob/main/laion-a... for more details.

What's remarkable is this: https://github.com/LAION-AI/laion-datasets/blob/main/laion-a...

That aesthetic predictor was apparently trained on only 4000 images. If my thinking is correct, imagine the impact those 4000 ratings have had on all of the output of this model.

You can see samples (some NSFW) of different images from the original training set in different rating buckets here, to get an idea of what was included or not in those training steps. http://3080.rom1504.fr/aesthetic/aesthetic_viz.html


That is really a shame, because all I really want is a version of Craiyon that I can modify and run on my own hardware.

The amount of enjoyment I have derived from playing with Craiyon over the last two months is ridiculous.


IIRC Craiyon runs Dalle-mega. https://huggingface.co/dalle-mini/dalle-mega

Note I think you need 16gb of VRAM to run it.


You can run craiyon / dalle-mini on a card with 8GB of VRAM if you decrease batch size to 1 and skip the CLIP step. Takes about 7 sec to generate an image on a 3070.

I started with https://github.com/borisdayma/dalle-mini/blob/main/tools/inf... and pared it down.


Have you checked out MidJourney? Makes Craiyon look like crayons :P


Craiyon is free, whereas Midjourney is not. If you want MJ level quality, check out Disco Diffusion or go straight to Visions of Chaos, which runs just about every AI diffusion script in existence. The dev is very active and adds new features every couple days, such as recently the ability to train your own diffusion models, which I've been doing the last 3 days nonstop on my little 3060 Ti (8GB VRAM, which is barely sufficient to run at mostly default settings).


MidJourney does give you 25 minutes of free compute time though. Which is enough for at least trying it ~40 times

I've checked out Disco Diffusion but hadn't heard of Visions of Chaos, thanks. The biggest shortcoming to DD is there's simply not yet a sufficiently trained model to produce stuff to the level of MidJourney or Craiyon


Are they giving out trials without an invite now? I was invited from someone who was paying, became addicted by the end of my trial, subscribed for a month, then gave out invites to friends, some of which ended up also paying - not a bad business model! It was bad timing though, the day after I joined, I discovered Disco Diffusion and haven't stopped rendering since (roughly 10k images rendered, mostly for animations). It takes longer, the results are often less realistic compared to Midjourney, Dall-E (1/2) or Stable Diffusion (which I've been toying with for a few weeks), but it's somehow much more satisfying having to wait xx minutes for a render to complete, running on your own local PC, not having to use bots with 1000 other people in a channel spamming their prompts, and having TONS of parameters to play around with. I have a google drive full of docs from my own studies, comparing parameter values, models, prompts, etc. I'm really looking forward to Stable Diffusion releasing their models, I know VoC will add those models as soon as they are available. On top of that, VoC has been adding support for diffusion models (of which I'm training my own), but there are new ones added constantly as more and more people build models for e.g. pixel art, medieval style, monochromatic, etc.

Also the results vary drastically if you have enough VRAM to load more models e.g. a 3090 (24GB) or an A6000 (48GB). I've been saving money and waiting impatiently for 4090s to drop. Check out the Disco Diffusion or VoC Discord - people post their works in there and often you will see results that make you wonder if they're cheating ;)


Which are the best models the 3090 enables you to load?


I'd start with VITL14 and add one or more RN types. I personally like to use multiple VIT and RN models just to fill up VRAM. What exactly will fit depends on your output resolution and requires a lot of trial and error. I always have Task Manager open to monitor VRAM usage. In general, VIT = more realistic, RN = more artistic. It can take a lot of experimentation to find what exactly tickles your fancy. I constantly change which models I'm using depending on what I'm going for. This redditor did a nice comparison[0], and there are many more "studies" for which models to use - you can google around, there are new articles/studies being posted daily.

You can also try disabling use_checkpoints if you have extra VRAM, since it will render a bit faster (but uses more VRAM since it doesn't save intermediary 'checkpoints' to disk).

When/if you get bored, try disabling use_secondary_models which will use a lot more VRAM but can deliver results on a completely different level. You will likely struggle for a few days figuring out which parameters to tweak to get good results (e.g. tv_scale, sat_scale, etc, which are otherwise AFAIK ignored).

In any case I recommend reading A Traveler's Guide to the Latent Space, which I call "The Bible" since it covers so many topics and has links to various studies and will keep you busy reading for months ;)

Also check out the Discord for Disco Diffusion and Visions of Chaos, as you can read endless tips and tricks to getting amazing results.

Have fun! :)

[0] https://www.reddit.com/r/DiscoDiffusion/comments/t7p4bi/seas...

[1] https://sweet-hall-e72.notion.site/A-Traveler-s-Guide-to-the...


I think I might disagree with your assessment of DD.

I can't use it, but check out this guy's work. incredible detail

https://instagram.com/textrnr




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: