Register and share your invite link to earn from video plays and referrals.

Search results for GaussianSplatting
GaussianSplatting community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including GaussianSplatting
From a single casual smartphone video, build a free-viewpoint "moving 3D human" — no studio, no camera rig required. Title: 4DAnyone: Create Anyone in 4D from a Casual Monocular Video URL: A framework that turns an uncalibrated monocular video into multi-view-consistent video, then into 4D Gaussian Splatting. Three highlights. 📦 Highlight 1: Compress reference context from O(N) to O(1) (RCP) 4D reconstruction needs dozens of views, but conditioning on generated views bloats the context and weakens guidance. Reference Context Packing packs references into fixed-length mixed-resolution context, keeping compute constant as views grow. 🔄 Highlight 2: Rotate view groups by noise level (TCR) Independently denoised view groups can't share info, so structure drifts. Target Context Routing cyclically shifts groups at high noise to propagate global structure, then fixes groups at low noise to stabilize detail. The switch at t/T=0.2 was optimal. 🎮 Highlight 3: Game-engine data for in-the-wild generalization 38k in-house multi-view videos (318 actors, 24 cameras) plus real light-stage and monocular data, trained via a 3-stage curriculum. It hits PSNR 24.33 on generation consistency, clearly beating prior methods. Photorealistic 4D avatars from phone video in ~40 min — free-viewpoint capture for everyone. #4DReconstruction# #GaussianSplatting#
Show more
Reconstruct a moving, deforming object in 4D (3D + time) from a single-camera video — even under heavy occlusion and large non-rigid motion 🎥 Title: Lift4D: Harmonizing Single-View 3D Estimation for 4D Reconstruction In-the-Wild URL: 🎥 Overview A test-time optimization framework combining single-view 3D prediction with deformable 3D Gaussian Splatting. It fuses learned geometry/appearance priors with the video to handle severe occlusion and non-rigid motion. ❓ Problem it solves ・Directly predicting 4D is limited by scarce training data and generalizes poorly ・Initializing 3D and refining with video alone breaks down on in-the-wild scenes with large deformation and occlusion 💡 Methodology Three components: ・Causal Latent Conditioning: aligns an image-to-3D model temporally without retraining, coupling adjacent latents at the ODE level, with t₀ trading off consistency vs per-frame fidelity ・Deformable Gaussians: sparse control nodes with time-varying SE(3) transforms + linear blend skinning deform a canonical representation across frames ・Occlusion-aware refinement: compares only visible pixels, while a view-conditioned diffusion prior completes unobserved surfaces 📊 Results ・Compared against STAG4D, PAD3R, L4GM, DreamMesh4D, V2M4, BANMo ・On Pexels data, point-tracking error EPE 0.072 (vs 0.119–0.211 for competitors) ・~30 min per 32-frame video on a single H200 GPU #3DReconstruction# #GaussianSplatting#
Show more
3D Gaussian Splatting for Autonomous Vehicle Simulation
Trains 3D Gaussian Splatting in 100 seconds
Introducing Hy3D WorldClaw——an agentic workflow that generates large scale 3D open worlds from text prompts.🚀🚀🚀 Not video, Not Gaussian Splatting, Every scene generated by WorldClaw is freely explorable and built entirely from editable, game-ready 3D assets with high-quality geometry and textures. Project Page: 
Show more
0
79
2.5K
260
Forward to community
Closing the sim-to-real gap for humanoid robotics starts with better simulation environments. Join us, @NianticSpatial, and @FlexionAI on Aug. 12 at 11 AM PT to see a real-to-sim workflow using 3D Gaussian splatting, Omniverse NuRec, Isaac, and OpenUSD. 📅 Add to calendar:
Show more
Spark 2.0 is here! 🚀 We’re redefining what’s possible on the web with a streamable LoD system for 3D Gaussian Splatting. Built on Three.js, you can now stream massive 100M+ splat worlds to any device from mobile to VR using WebGL2. All open-source. Dive into the tech 👇
Show more
0
41
2.1K
323
Forward to community
🪑 Insert an object into an image while specifying its exact 3D orientation and position. DIRECT solves the 3D-pose control that text leaves ambiguous and parameters struggle with, by decomposing visual proxies. Title: Direct 3D-Aware Object Insertion via Decomposed Visual Proxies URL: 📝 Overview DIRECT is a diffusion-based method that inserts a reference object into an image with explicit control over its 3D pose and position. It decomposes the insertion condition into geometry, appearance, and context, injected through independent pathways. ❓ Challenges Solved Existing insertion methods formulate the task as 2D inpainting and can't control 3D pose. Text guidance is spatially ambiguous, and parametric 3D methods can't translate abstract parameters into correct geometric projections. 💡 Methodology & Proposed Approach ・A user-manipulated 3D proxy rendered at the target pose provides geometry guidance ・Appearance (the reference's high-fidelity look) and context (background semantics) are injected independently via separate LoRA adapters and positional embeddings to avoid feature entanglement ・TRELLIS lifts the image into a coarse 3D shape, refined with VGGT and 3D Gaussian Splatting ・Built on FLUX.1-Fill, it uses shape-decomposed mask augmentation and progressive-resolution training to avoid overfitting 🎯 Use Cases It fits virtual staging, e-commerce product photography, creative work needing precise spatial control, and photorealistic AR/VR content generation. 📊 Experimental Results ・On the FLUX backbone it reaches PSNR 23.09, LPIPS 0.147, and matching error 17.8, beating baselines on all metrics ・It stays stable across large 0-180 degree pose changes and preserves fine details even under 3D-reconstruction degradation ・Hybrid-data training raised CLIP-I from 0.904 to 0.943 ・For symmetric object orientation, RGB geometry guidance outperformed normal maps #3DGeneration# #ImageEditing#
Show more
Most people see a street. He sees $300-600 per block. A 24-year-old from Chengdu figured out that every hotel, every apartment, every commercial space within walking distance is an untapped asset. One nobody has packaged yet. He straps a rig to his back, walks in, spends twenty minutes scanning the space, and leaves with a file that lets anyone on earth stand inside that room from their couch. The client pastes a link on their booking page. Guests tour the property before they arrive. Cancellations drop. Reviews go up. He gets paid $400 for the scan. $99 every month for hosting. The technology: 3D Gaussian Splatting. Free on GitHub since 2023. The app: Luma AI. Also free. The page he delivers: built by Claude in ten minutes. Total tool cost: $20/month. Month one: $3,500. Month six: $18,000. The streets haven't changed. He just started charging for them.
Show more
SOMEONE JUST KILLED THE REAL ESTATE INDUSTRY A guy scanned an entire house with his phone. Uploaded it. Now anyone on Earth can walk through it in a browser tab. No app. No VR. No agent. No appointment. Click → you’re inside. Every room. Every angle. Every shadow. Photoreal. The numbers are insane: - Agent fee on a $500k home: $15,000 - Cost to make this scan: ~$200 - Time to “tour” 50 houses: one evening - File size: smaller than a TikTok The science is wild too: It’s called 3D Gaussian Splatting instead of polygons (how games render), it uses millions of tiny glowing “splats” of color and depth. AI reconstructs reality from your photos. The result loads on a phone and looks like you’re THERE. The grift opportunity is even wilder: Freelancers are already charging $300–$800 per scan for realtors, Airbnbs, venues, car dealers, museums. One person + one phone + one weekend = a business. Open source. Built on PlayCanvas. Free GitHub:
Show more
Claude for Excel, PowerPoint, and Word are now generally available, and Claude for Outlook is in public beta. As Claude moves between your Microsoft apps, it carries the full context of your conversation.
Show more
0
1.1K
34K
2.8K
Forward to community