Register and share your invite link to earn from video plays and referrals.

Search results for ImageEditing
ImageEditing community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including ImageEditing
🖼 Test-time scaling for image editing tends to hand every edit the same compute budget, wasting a lot of it. By allocating budget by difficulty and pruning with edit-specific verification, this work hits up to 2.2x speedup while preserving quality. Title: From Scale to Speed: Adaptive Test-Time Scaling for Image Editing URL: 📝 Overview ADE-CoT is a test-time scaling method tailored to goal-directed image editing. Instead of reusing Image-CoT methods built for text-to-image generation, it combines three strategies, difficulty-aware allocation, edit-specific early verification, and opportunistic stopping, to cut compute substantially while preserving quality. ❓ Challenges Solved Prior methods had three mismatches. ・Fixed sampling budgets waste compute on easy edits that barely improve ・General MLLM scores wrongly prune about 40% of samples that start low but ultimately score high ・Large-scale sampling produces redundant identical correct outputs, adding needless compute 💡 Methodology & Proposed Approach ・It reads edit difficulty, giving easy edits a minimal budget and expanding the search for hard ones ・A one-step preview estimates clean latents from noisy intermediates without extra denoising, making early verification reliable ・Grounded SAM2 checks that only the intended region changed, and DINOv2 embeddings remove redundant candidates ・It generates candidates sequentially and stops, via depth-first opportunistic stopping, once enough intent-aligned results are found 🎯 Use Cases It fits complex pose changes, multi-object removal or replacement, fine-grained regional edits, multi-turn editing, and high-quality editing under compute constraints, and is especially valuable where inference cost matters, like a production image-editing API. 📊 Experimental Results ・On GEdit-Bench, FLUX.1 Kontext is 2.2x, BAGEL 1.8x, and Step1X-Edit 2.0x faster than Best-of-N ・Reasoning efficiency more than doubles on a fixed 32-sample budget, and outcome efficiency rises 4.9x, 2.7x, and 2.9x across three benchmarks ・On hard multi-object edits like "remove the person standing next to the lady in white," it fixes the baseline's misidentification #ImageEditing# #DiffusionModels#
Show more
🪑 Insert an object into an image while specifying its exact 3D orientation and position. DIRECT solves the 3D-pose control that text leaves ambiguous and parameters struggle with, by decomposing visual proxies. Title: Direct 3D-Aware Object Insertion via Decomposed Visual Proxies URL: 📝 Overview DIRECT is a diffusion-based method that inserts a reference object into an image with explicit control over its 3D pose and position. It decomposes the insertion condition into geometry, appearance, and context, injected through independent pathways. ❓ Challenges Solved Existing insertion methods formulate the task as 2D inpainting and can't control 3D pose. Text guidance is spatially ambiguous, and parametric 3D methods can't translate abstract parameters into correct geometric projections. 💡 Methodology & Proposed Approach ・A user-manipulated 3D proxy rendered at the target pose provides geometry guidance ・Appearance (the reference's high-fidelity look) and context (background semantics) are injected independently via separate LoRA adapters and positional embeddings to avoid feature entanglement ・TRELLIS lifts the image into a coarse 3D shape, refined with VGGT and 3D Gaussian Splatting ・Built on FLUX.1-Fill, it uses shape-decomposed mask augmentation and progressive-resolution training to avoid overfitting 🎯 Use Cases It fits virtual staging, e-commerce product photography, creative work needing precise spatial control, and photorealistic AR/VR content generation. 📊 Experimental Results ・On the FLUX backbone it reaches PSNR 23.09, LPIPS 0.147, and matching error 17.8, beating baselines on all metrics ・It stays stable across large 0-180 degree pose changes and preserves fine details even under 3D-reconstruction degradation ・Hybrid-data training raised CLIP-I from 0.904 to 0.943 ・For symmetric object orientation, RGB geometry guidance outperformed normal maps #3DGeneration# #ImageEditing#
Show more
Qwen-Image-2.1-viggle-turbo v0.2 is out on Hugging Face Text-to-image and image editing in 6 steps about 5× faster than the 40-step Qwen-Image-2.1 model:
Show more
Grok Image 2.0 is now ranked #2# globally on Arena for both text-to-image generation and image editing
BREAKING: Grok Imagine Image 2.0 just climbed from #18# to #4# on Artificial Analysis’ Text-to-Image Leaderboard. 🚀 With a 1,154 Elo rating, it’s now the highest-ranked model outside OpenAI and has landed on the Pareto frontier for quality vs. price. • #4# in Text-to-Image, up 14 places • #10# in Image Editing, up from #16# • Supports text-to-image, image editing and up to 5-image references • Biggest gains in knowledge, text rendering and lighting Grok Imagine Image 2.0 is available on Grok, the xAI API, fal and Replicate. A massive jump in a single generation. ⚡
Show more
Marigold V2 turns an image-editing DiT into a single-step model for sharp, detailed dense prediction.📜 Apache 2.0. 🤖 📄 🏆 Best zero-shot results across all evaluated depth datasets among models trained on comparable data. AbsRel improves by 16%–26% over the previous best on KITTI and ETH3D. 🔍 Fur, foliage, fine wires, and object boundaries stay crisp. The two-stage iREPA and SinkLoss recipe tackles the smoothing and flying-pixel artifacts common in diffusion-based depth models. 🧩 The same framework reaches SOTA results in depth completion, see-through depth, surface normals, and intrinsic image decomposition. Depth completion records the lowest RMSE across all four reported benchmarks. ⚡ A pretrained Qwen-Image-Edit DiT becomes a single-step dense predictor through lightweight adaptation. Training takes less than a week on one 32 GB GPU.
Show more
Grok Imagine’s image editing literally blows my mind I just edited this entire image from scratch using Grok Imagine and its full set of editing tools The segment + precise editing tools are insanely useful Instead of regenerating the entire image, you can target exactly what you want to change while keeping everything else intact That level of control makes editing much faster and more practical Grok Imagine is becoming a seriously powerful image editor….not just an image generator
Show more
Grok Imagine Image 2.0 just got one of the coolest image editing features yet Hover over any part of your image, select the segment, and edit it instantly This is seriously one of the smartest and most intuitive AI image editing features I’ve seen
Show more
Grok Imagine Image 2.0 just released and it's a MASSIVE upgrade Image 2.0 brings: • Much better instruction following • Sharper text + significantly better typography • Complex layouts that actually hold together • Precise region editing with the Magic Wand • Segmentation to edit specific parts of an image • One-click background removal • Multi-reference editing with up to 5 images at once • Smart Resize that expands an image into almost any aspect ratio • Much stronger consistency across generations and edits • New ready-made templates for product shots, headshots, e-commerce, game assets, icons, merch and more And the performance is already at the top: Grok Image 2.0 now ranks #2# in the world for BOTH text-to-image generation and image editing on Arena
Show more
Alibaba 😃can release a new open image model - Swift-Image 6B a compact unified model (lighter than Flux 2 klein) for: - text-to-image generation -single-image editing, and -multi image editing. paper:👇
Show more