Register and share your invite link to earn from video plays and referrals.

Search results for 4D生成
4D生成 community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including 4D生成
Cinema 4D. Blender. Houdini. Standalone. Different interfaces, same underlying Octane. Day 3 of Octane Foundations looks beyond the DCC to the engine underneath.
A NATO RQ-4D Phoenix drone has conducted a night reconaissance flight over the Black Sea. The ISR drone is currently landing at Sigonella Air Base after flying for several hours.
Show more
When generating 4D worlds from video models, you never had to detour through RGB — a new method goes directly from latent space to 4D. Beyond Pixels: From Video Priors to 4D Worlds 🔍 Overview The conventional cascade — video generation → RGB decoding → 4D reconstruction — inherently loses geometric information at the RGB step and amplifies generator-specific artifacts. This work proposes "latent-to-4D generation," directly connecting final denoised video VAE latents to a 4D decoder without RGB decoding. 🛠 Problems Solved · Geometric information degrades during RGB decoding · Generator-specific noise propagates through RGB into 4D reconstruction · New generators require retraining the entire 4D module All three are bypassed by treating the VAE latent space as the direct interface to 4D synthesis. ⚙ Method Three components: · Alignment Module (𝒜ϕ): Trilinear resampling + 3D convolutions transform video latents into the 4D decoder's token space · L4AR Attention: Hierarchical frame-wise spatial alignment + global temporal attention across frames · 4D Decoder: Jointly predicts per-frame camera parameters (9D) and dense world-space point maps Only these three components are trained (rank-16 LoRA included). Video generators, VAE, and Transformers stay frozen, trained on ~1K annotated clips. 📊 Results DINO-F1 (text-to-4D): 57.01–57.09 (Ours) vs. 53.27–54.21 (cascade baseline). Image-to-4D: 61.60 vs. 55.79 — a 5.81-point improvement. Human evaluators preferred our geometry and completeness in 66.8–72.1% of comparisons. Robustness test at contamination ρ=0.6: point-map drift 0.0053 vs. 0.3827 baseline (71× more stable). 🌐 Generalization A single checkpoint works unchanged across Wan T2V at 14B and 1.3B, and Wan I2V — all sharing the same VAE. Motion, appearance, camera, and trajectory controls remain functional without any retraining. #4DGeneration# #VideoAI#
Show more
Imagine some 4D baddie doing this to your balls.
0
105
867
45
Forward to community
The chessboard is 4D. ♟️ Xi: “Elon, I heard you named your social media and AI companies after me. What can I do for you?” Elon: “Approve FSD.”
Generates realistic videos with explicit 4D geometric control
Dive into the weekend 🌊 Ripplor - fully procedural water for Cinema 4D. Ripples, wakes, and rain that lands with real crowns and jets. No sim cache. Any renderer. Drop it on a mesh. Press play. $10 off with code SHARK until Sep 1
Show more
"Depth Anything V4" Depth Anything V4 reconstructs dynamic 4D scenes from monocular video using Riemannian Flow Matching over 4D Gaussian parameters. Instead of forcing rotations, scales, and opacity through Euclidean space, it follows their natural manifolds, cutting invalid Gaussians from 12.3% to 0.01% and improving F1 from 0.762 to 0.806 over an otherwise identical MLP baseline. Moving from frame-wise depth prediction to probabilistic 4D scene understanding.
Show more
Reconstruct a moving, deforming object in 4D (3D + time) from a single-camera video — even under heavy occlusion and large non-rigid motion 🎥 Title: Lift4D: Harmonizing Single-View 3D Estimation for 4D Reconstruction In-the-Wild URL: 🎥 Overview A test-time optimization framework combining single-view 3D prediction with deformable 3D Gaussian Splatting. It fuses learned geometry/appearance priors with the video to handle severe occlusion and non-rigid motion. ❓ Problem it solves ・Directly predicting 4D is limited by scarce training data and generalizes poorly ・Initializing 3D and refining with video alone breaks down on in-the-wild scenes with large deformation and occlusion 💡 Methodology Three components: ・Causal Latent Conditioning: aligns an image-to-3D model temporally without retraining, coupling adjacent latents at the ODE level, with t₀ trading off consistency vs per-frame fidelity ・Deformable Gaussians: sparse control nodes with time-varying SE(3) transforms + linear blend skinning deform a canonical representation across frames ・Occlusion-aware refinement: compares only visible pixels, while a view-conditioned diffusion prior completes unobserved surfaces 📊 Results ・Compared against STAG4D, PAD3R, L4GM, DreamMesh4D, V2M4, BANMo ・On Pexels data, point-tracking error EPE 0.072 (vs 0.119–0.211 for competitors) ・~30 min per 32-frame video on a single H200 GPU #3DReconstruction# #GaussianSplatting#
Show more