Register and share your invite link to earn from video plays and referrals.

Search results for 3DReconstruction
3DReconstruction community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including 3DReconstruction
Reconstruct a moving, deforming object in 4D (3D + time) from a single-camera video — even under heavy occlusion and large non-rigid motion 🎥 Title: Lift4D: Harmonizing Single-View 3D Estimation for 4D Reconstruction In-the-Wild URL: 🎥 Overview A test-time optimization framework combining single-view 3D prediction with deformable 3D Gaussian Splatting. It fuses learned geometry/appearance priors with the video to handle severe occlusion and non-rigid motion. ❓ Problem it solves ・Directly predicting 4D is limited by scarce training data and generalizes poorly ・Initializing 3D and refining with video alone breaks down on in-the-wild scenes with large deformation and occlusion 💡 Methodology Three components: ・Causal Latent Conditioning: aligns an image-to-3D model temporally without retraining, coupling adjacent latents at the ODE level, with t₀ trading off consistency vs per-frame fidelity ・Deformable Gaussians: sparse control nodes with time-varying SE(3) transforms + linear blend skinning deform a canonical representation across frames ・Occlusion-aware refinement: compares only visible pixels, while a view-conditioned diffusion prior completes unobserved surfaces 📊 Results ・Compared against STAG4D, PAD3R, L4GM, DreamMesh4D, V2M4, BANMo ・On Pexels data, point-tracking error EPE 0.072 (vs 0.119–0.211 for competitors) ・~30 min per 32-frame video on a single H200 GPU #3DReconstruction# #GaussianSplatting#
Show more
This 3D reconstruction shows the immense amount of depth information a Tesla can collect from just a few seconds of video from the vehicle's 8 cameras
0
463
22.9K
3.2K
Forward to community
🪑 Insert an object into an image while specifying its exact 3D orientation and position. DIRECT solves the 3D-pose control that text leaves ambiguous and parameters struggle with, by decomposing visual proxies. Title: Direct 3D-Aware Object Insertion via Decomposed Visual Proxies URL: 📝 Overview DIRECT is a diffusion-based method that inserts a reference object into an image with explicit control over its 3D pose and position. It decomposes the insertion condition into geometry, appearance, and context, injected through independent pathways. ❓ Challenges Solved Existing insertion methods formulate the task as 2D inpainting and can't control 3D pose. Text guidance is spatially ambiguous, and parametric 3D methods can't translate abstract parameters into correct geometric projections. 💡 Methodology & Proposed Approach ・A user-manipulated 3D proxy rendered at the target pose provides geometry guidance ・Appearance (the reference's high-fidelity look) and context (background semantics) are injected independently via separate LoRA adapters and positional embeddings to avoid feature entanglement ・TRELLIS lifts the image into a coarse 3D shape, refined with VGGT and 3D Gaussian Splatting ・Built on FLUX.1-Fill, it uses shape-decomposed mask augmentation and progressive-resolution training to avoid overfitting 🎯 Use Cases It fits virtual staging, e-commerce product photography, creative work needing precise spatial control, and photorealistic AR/VR content generation. 📊 Experimental Results ・On the FLUX backbone it reaches PSNR 23.09, LPIPS 0.147, and matching error 17.8, beating baselines on all metrics ・It stays stable across large 0-180 degree pose changes and preserves fine details even under 3D-reconstruction degradation ・Hybrid-data training raised CLIP-I from 0.904 to 0.943 ・For symmetric object orientation, RGB geometry guidance outperformed normal maps #3DGeneration# #ImageEditing#
Show more
asked GPT-6 Astra to explain the recent OpenAI–Hugging Face incident as a cinematic 3d reconstruction effort: max took 44 minutes 6.73m tokens used 15% of my weekly usage
0
53
1.2K
89
Forward to community
The future of football replay is more immersive ⚽ Lenovo AI-powered 3D Digital Avatars use advanced 3D reconstruction for ultra-realistic player views at FIFA World Cup 2026™.
Show more
That grainy tangle of colors is real brain tissue, mapped in more detail than any human brain sample ever has been. The sample came from a woman's temporal lobe, removed during surgery to treat her epilepsy, and donated for research afterward. A team from Google Research and Jeff Lichtman's lab at Harvard sliced it into roughly 5,000 wafer-thin sections, imaged each one with an electron microscope, then used machine learning to digitally stitch everything back into one continuous 3D reconstruction. No human could trace this by hand, the resulting dataset alone runs to 1.4 petabytes, more storage than most people will use in a lifetime. Inside that single cubic millimeter, roughly one millionth of an entire human brain, the team counted about 57,000 individual cells, nearly 150 million synapses connecting them, and 230 millimeters of blood vessels threading through the tissue. The colors in images like this one aren't decorative, they're how researchers visually separate individual neurons and cell types from the tangle around them so they can trace each one's path. The map turned up some genuine surprises too. Most neuron pairs connect through just a couple of synapses, but researchers found rare pairs linked by up to 50 connections at once. They also spotted neurons whose branches curl around and knot into themselves, something nobody had documented before, along with pairs of neurons that turned out to be near perfect mirror images of each other. Knowing this level of detail exists for just one millionth of a brain, how far off do you think a full human connectome actually is?
Show more
DAILY AI BRIEF 🗞️ — Sept 2 OPENAI 🔥: > Official “Path to Astra” post is out. Astra is the first model OpenAI has designated as Critical for cybersecurity under the Preparedness Framework. It scored 100% on ExploitBench, found 2 zero-days in evals, and is “coming soon” — advanced cyber tools stay limited to testers / Daybreak Blue at first. > @M1Astra reported on a fresh Astra test prep the same day. Models in testing: vega-alpha (new) and ultima-alpha. GOOGLE 🔥: > Gemini 3.8 Flash is already answering on Gemini and in the Gemini app. Some people still have 3.7 selected and get 3.8 anyway. > WSJ: Google engineers preferred it to Opus for coding in Jetski tests. Official drop looks like today. > Agentic video understanding is on Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. The model hunts transcript/audio/frames instead of eating the whole file. Up to 88% fewer tokens, 66% lower cost, ~7% better accuracy on long video. API + AI Studio now, Gemini app later. ANTHROPIC 🔥: > Claude Fable 5.1 (and Mythos 5.1) is live. 52.6% on Terminal-Bench-Science 0.1 (more than 2x Fable 5) and 55.8% vs 42.0% on Terminal-Bench 4.0. Same list price as Fable 5, 75% cheaper cache reads — about 25% cheaper typically, up to 45% on heavy agent runs. META 🔥: > Muse Voice Transcribe is live — MSL’s first real-time audio perception model. SOTA streaming speech-to-text, native diarization (20+ speakers) and endpointing in one model, multilingual with code-switching. Rolling out on the Meta Model API, Meta AI for Mac, and Muse Code. Zero-data-retention tier available. XAI 🔥: > Elon: “Grok 4.7 comes out in 10 days.” That’s ~Sept 12. Reply to Tobi on Grok 4.6. ALIBABA 🔥: > Qwen3.8-Max-0902 is live on QwenCloud. Same 2.4T / 1M-context Max line, with extra post-training on coding and cowork. $2 in / $6 out per 1M tokens. Arena has it #1# on Code Arena: WebDev at 1691 — 3 pts above Claude Opus 5 (Max) and 22 pts above the previous 3.8-Max. WORLD LABS 🔥: > Fei-Fei Li’s lab shipped Atlas, an omni world model. Few photos > pixel-perfect camera control, up to 1 min of 1440p video, plus 3D reconstruction (point clouds / Gaussian splats). Early access only; it will power future Marble. No paper, no public API yet. * Too much is happening, and I have some scoops planned for today, so I don't want to spam the algo. ** I used Grok to compose this brief, cherry-picking the news and doing some post-editing.
Show more
The 2023 Holiday Update rolls out next week Here’s what’s coming... Custom Lock Sounds Replace the horn lock sound of your vehicle with another sound—like a screaming goat 🐐 LAN Party on Wheels Play your favorite games on the rear touchscreen 🎮 Rear Screen Bluetooth Headsets Parents, rejoice. Rear passengers can now use wireless Bluetooth headphones when watching shows or playing games on the rear screen Apple Podcasts Listen to millions of the world’s most popular podcasts @ApplePodcasts Tesla App Trip Planner Use the Tesla mobile app to plan a multi-stop trip and send it to your vehicle Speed Cameras on Route Navigation now includes speed cameras, stop signs & traffic lights Automatic 911 Calls Your vehicle will automatically call 911 if an accident triggers the airbags Blind Spot Indicators The blind spot camera will alert you with red shading when your turn signal is on and a car is detected in your blind spot More Live Sentry Cameras When viewing vehicle surroundings from the Tesla app, you'll now have access to the left and right pillar cameras for a total of 7 angles Light Show Enjoy a thunderous new Light Show called 'The Arrival' Castle Doombad Castle Doombad is now available in the Tesla Arcade—plus updates to Beach Buggy, Polytopia & Vampire Survivors High Fidelity Park Assist See a 3D reconstruction of your surroundings while parking *Availability varies by model and location
Show more
0
1.3K
16.2K
2.3K
Forward to community
From a single casual smartphone video, build a free-viewpoint "moving 3D human" — no studio, no camera rig required. Title: 4DAnyone: Create Anyone in 4D from a Casual Monocular Video URL: A framework that turns an uncalibrated monocular video into multi-view-consistent video, then into 4D Gaussian Splatting. Three highlights. 📦 Highlight 1: Compress reference context from O(N) to O(1) (RCP) 4D reconstruction needs dozens of views, but conditioning on generated views bloats the context and weakens guidance. Reference Context Packing packs references into fixed-length mixed-resolution context, keeping compute constant as views grow. 🔄 Highlight 2: Rotate view groups by noise level (TCR) Independently denoised view groups can't share info, so structure drifts. Target Context Routing cyclically shifts groups at high noise to propagate global structure, then fixes groups at low noise to stabilize detail. The switch at t/T=0.2 was optimal. 🎮 Highlight 3: Game-engine data for in-the-wild generalization 38k in-house multi-view videos (318 actors, 24 cameras) plus real light-stage and monocular data, trained via a 3-stage curriculum. It hits PSNR 24.33 on generation consistency, clearly beating prior methods. Photorealistic 4D avatars from phone video in ~40 min — free-viewpoint capture for everyone. #4DReconstruction# #GaussianSplatting#
Show more