Feed-forward 3D Gaussian Splatting (3DGS) enables efficient and
generalizable 3D reconstruction, but current feed-forward 3DGS
methods for scene understanding remain largely category-oriented.
In contrast, instance-aware 3DGS methods typically rely on
per-scene optimization and often decouple reconstruction from
instance and semantic learning, limiting reciprocal interactions
among them.
We present InstanceSplat, a unified feed-forward 3DGS framework
for generalizable 3D reconstruction and instance-aware scene
understanding from pose-free multi-view images. In a single
forward pass, InstanceSplat constructs an instance-aware Gaussian
representation that jointly encodes appearance, geometry, instance
identity, and language-aligned semantics. Shared 3D Gaussians
ground instance identities across views, producing renderable and
cross-view-consistent instance features.
To allow reconstruction and scene understanding to benefit from
each other, we further design an instance-centric learning strategy
that connects reconstruction, instance learning, and semantic
learning through shared instance structure. Specifically, instance
cues guide reconstruction, language-aligned semantics strengthen
the discrimination of confusing same-category instances, and
instance regions aggregate semantic evidence into coherent
object-level predictions.