Detailed Information

Cited 0 time in webofscience Cited 0 time in scopus
Metadata Downloads

SAGE: Segmentation-Aware 3D object extraction from single images

Authors
Jeong, JuyongKwon, SungrokLee, HajeongPark, Jong-Il
Issue Date
Feb-2026
Publisher
SPIE
Keywords
3D Reconstruction; Indoor Scene Understanding; Object Extraction; Point Cloud Processing; Semantic Segmentation
Citation
Proceedings of SPIE - The International Society for Optical Engineering, v.14072
Indexed
SCOPUS
Journal Title
Proceedings of SPIE - The International Society for Optical Engineering
Volume
14072
URI
https://scholarworks.bwise.kr/hanyang/handle/2021.sw.hanyang/219725
DOI
10.1117/12.3102254
ISSN
0277-786X
1996-756X
Abstract
Recent progress in vision-based 3D reconstruction has enabled dense point cloud generation directly from a single RGB image, but most existing methods provide only geometric information without semantic context. This limitation hinders object-level understanding and constrains downstream applications such as scene analysis and augmented reality. To address this limitation, we propose a segmentation-aware 3D object extraction framework that combines VGGT, a state-of-to-art geometry transformer, with SegFormer, an efficient semantic segmentation to assign pixel-level category labels while VGGT reconstructs a dense 3D point cloud from the same image. The segmentation results are projected onto the reconstructed points, producing a labeled 3D point cloud where each point is enriched with both geometric and semantic information. Using this representation, we perform clustering within each label by considering point count and density, enabling the segmentation. This approach enables object-level separation directly from single images, allowing labeled 3D reconstructions to be exported as GLB files for visualization and further analysis. Experiments conducted on multiple indoor scenes demonstrate that our system successfully reconstructs point clouds with semantic labels and separates objects into clusters. By unifying semantic segmentation with geometric reconstruction, we propose a robust framework for semantic 3D modeling and object-aware processing.
Files in This Item
Go to Link
Appears in
Collections
서울 공과대학 > 서울 컴퓨터소프트웨어학부 > 1. Journal Articles

qrcode

Items in ScholarWorks are protected by copyright, with all rights reserved, unless otherwise indicated.

Related Researcher

Researcher Park, Jong-Il photo

Park, Jong-Il
COLLEGE OF ENGINEERING (SCHOOL OF COMPUTER SCIENCE)
Read more

Altmetrics

Total Views & Downloads

BROWSE