One of the things they seem to be emphasizing here is the UX around being able to place specific elements where you want them in an image. If the positions of the components in the overall composition are very important, this seems to make that a lot easier and kind of reminds me of InvokeAI.
Ideogram V4, an open-weight model released back in June can also do this [1], but you have to use a relatively cumbersome JSON structure to describe all the different bounding boxes. So it’s definitely a bit of a hassle.
I'll probably be waiting until it goes open-weight (hopefully soon) like they did with Flux.2 / Klein.
I always hated Comfy's node based UI, but agents make it tolerable. Now I just have them set up a workflow and I go in and tweak it manually if the results aren't where I want them. I even have agents cherry doing multiple runs and cherry picking the best outputs, models have gotten good enough that it's a real time saver, assuming you have references they can and a rubric to check against.
Yeah, given how much better Ideogram v4 outputs are when you use the proper structured JSON (background, elements, etc.) I think most users probably have stuck some kind of Qwen/Gemma-based LLM between their raw prompt and the CLIP encoder.
Does anyone know if it can be used to generate accurate frame-by-frame sprite sequences? I found that no image model can do this well (with sufficient fidelity) - neither with one shot (full spritesheet), nor single frame conditioning. It would be great if an imagegen model could do this. What I do now (I use my own tool https://github.com/acatovic/ai-game-studio) is basically generate a reference image, then condition on that image to generate a very short video, then extract and prune frames. Then I get indie-level sprite fidelity about 90% of the time.
The task you're describing is a video model task, not an image model task. It's inherently temporal.
Generate a sprite in an image editor, then use a video model to make the loop you want; then turn the resulting video back into individual sprite images.
So much negativity as usual and so little talk about the product, this is pretty impressive, well done, it seems to be filling decently a gap that everyone that has worked enough generating images with AI has faced.
I shouldn't have to scroll all the way down, click to use the thing, then fumble around to figure out what this does. It should be clear at the top of the first page (so I know right away that I don't need this).
Unless the version 3 announcement is your onboarding point, like it is for me just now. I can go research FLUX from here, but the original poster's point still holds.
I like the interface; very useful for some use cases that would otherwise be quite frustrating. Dislike that it's yet another platform held back by arbitrary moderation. You can't make a bicycle for the mind that locks if you try to ride it in the wrong direction.
> FLUX 3 Image is available under a commercial weights license for companies running image generation at scale. Fine-tune and deploy it on your own infrastructure. Reach out to us to learn more.
Agreed with other comments about the UX. I'm more interested in that than the model itself. Would like to start seeing UI like this where you get to choose the model and compare different models. Can't jump all over the internet to each model developers sandbox just to test their models. Doing it from one place would be nice.
That sort of steering ability that has been possible with the latest Gemini releases has been nice to work with over previous generations. It’s great to see this improve on the platform with declarative controls built into the API and coming soon as an open model.
One of the things they seem to be emphasizing here is the UX around being able to place specific elements where you want them in an image. If the positions of the components in the overall composition are very important, this seems to make that a lot easier and kind of reminds me of InvokeAI.
Ideogram V4, an open-weight model released back in June can also do this [1], but you have to use a relatively cumbersome JSON structure to describe all the different bounding boxes. So it’s definitely a bit of a hassle.
I'll probably be waiting until it goes open-weight (hopefully soon) like they did with Flux.2 / Klein.
[1] - https://docs.ideogram.ai/using-ideogram/getting-started/prom...
You just get an LLM to do the bounding box stuff or use the ComfyUI node that provides a GUI for bounding box generation
I always hated Comfy's node based UI, but agents make it tolerable. Now I just have them set up a workflow and I go in and tweak it manually if the results aren't where I want them. I even have agents cherry doing multiple runs and cherry picking the best outputs, models have gotten good enough that it's a real time saver, assuming you have references they can and a rubric to check against.
Yeah, given how much better Ideogram v4 outputs are when you use the proper structured JSON (background, elements, etc.) I think most users probably have stuck some kind of Qwen/Gemma-based LLM between their raw prompt and the CLIP encoder.
Ideogram 4.5 released 2 days ago with more features along these lines
https://www.youtube.com/watch?v=2mecWZgbaEg
Does anyone know if it can be used to generate accurate frame-by-frame sprite sequences? I found that no image model can do this well (with sufficient fidelity) - neither with one shot (full spritesheet), nor single frame conditioning. It would be great if an imagegen model could do this. What I do now (I use my own tool https://github.com/acatovic/ai-game-studio) is basically generate a reference image, then condition on that image to generate a very short video, then extract and prune frames. Then I get indie-level sprite fidelity about 90% of the time.
The task you're describing is a video model task, not an image model task. It's inherently temporal.
Generate a sprite in an image editor, then use a video model to make the loop you want; then turn the resulting video back into individual sprite images.
The UX looks amazing and very steerable, congrats to the team for focusing on the interface.
Chats can be awful user interfaces.
So much negativity as usual and so little talk about the product, this is pretty impressive, well done, it seems to be filling decently a gap that everyone that has worked enough generating images with AI has faced.
I shouldn't have to scroll all the way down, click to use the thing, then fumble around to figure out what this does. It should be clear at the top of the first page (so I know right away that I don't need this).
It is a version 3 of a product.If you actually used or needed these, you would have known right away what this is and how it works.
Unless the version 3 announcement is your onboarding point, like it is for me just now. I can go research FLUX from here, but the original poster's point still holds.
Yes, you can do that with AI now.
I like the interface; very useful for some use cases that would otherwise be quite frustrating. Dislike that it's yet another platform held back by arbitrary moderation. You can't make a bicycle for the mind that locks if you try to ride it in the wrong direction.
I think we are all waiting for the open weights or local model releases.
Has that been announced?
The website mentions:
> FLUX 3 Image is available under a commercial weights license for companies running image generation at scale. Fine-tune and deploy it on your own infrastructure. Reach out to us to learn more.
I guess the open ones would be non-commercial?
If it is anything like the previous release Flux.2 [dev] - then yeah it'll probably be a non-commercial license.
https://bfl.ai/legal/non-commercial-license-terms
Agreed with other comments about the UX. I'm more interested in that than the model itself. Would like to start seeing UI like this where you get to choose the model and compare different models. Can't jump all over the internet to each model developers sandbox just to test their models. Doing it from one place would be nice.
It's this a new model or a new ui?
https://news.ycombinator.com/item?id=49031796
What's new from the last post? GA?
The original article mentions:
> We will open up an early access phase for FLUX 3 Image in the following weeks.
Not sure if there was a separate post for early access or if they just skipped to this.
That sort of steering ability that has been possible with the latest Gemini releases has been nice to work with over previous generations. It’s great to see this improve on the platform with declarative controls built into the API and coming soon as an open model.
looks cool, eternally greatful these models are marked for open weight releases. Pretty excited
Unrelated to the flux images used for floppy disk archiving? Sigh.
Every new piece of software needs a name.
Flux is in the top 9000 of the most common words.