WithEveryone Unveils Unified Planning and Identity Grounding for Group Image Generation
cohere training
| Source: HF Papers | Original article
A new model called WithEveryone enables unified planning and identity grounding to reliably generate group images with multiple specified individuals.
A new AI model called **WithEveryone** promises to make group‑photo generation far more reliable. The system, released this week, can synthesize coherent images that include five to ten distinct reference identities while keeping each person recognisable. Unlike earlier generators that stumble when asked to place several known faces together, WithEveryone first decides which references belong in the scene, then creates a detailed plan that binds each identity to a specific location, pose and region before rendering the final picture.
The advance matters because identity‑preserving generation has long been a weak spot for text‑to‑image tools. When a prompt calls for multiple known individuals, models often mix up faces or collapse them into generic figures. By integrating planning and identity grounding into a single network, WithEveryone tackles the correspondence problem at its source, offering a more predictable workflow for advertisers, content creators and journalists who need to depict real people together without manual compositing.
The research team has open‑sourced the code on GitHub, inviting the community to test the approach on broader scenarios. Observers will watch for benchmark results that compare WithEveryone’s fidelity and consistency against existing engines such as Gemini or OpenAI’s latest image models. Further developments may include scaling the method to larger crowds, extending it to video, or coupling it with downstream tools for watermark removal or animation. As the field pushes toward more controllable, multi‑entity generation, WithEveryone marks a concrete step toward reliably “bringing everyone into the frame.”
Sources
Back to AIPULSEN