<?xml version="1.1" encoding="utf-8"?>
<article xsi:noNamespaceSchemaLocation="http://jats.nlm.nih.gov/publishing/1.1/xsd/JATS-journalpublishing1-mathml3.xsd" dtd-version="1.1" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"><front><journal-meta><journal-id journal-id-type="publisher-id">TACS</journal-id><journal-title-group><journal-title>Technology and Application of Computer Science</journal-title></journal-title-group><issn>2998-8926</issn><eissn>2998-8934</eissn><publisher><publisher-name>Art and Technology</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.61369/TACS.2026050048</article-id><article-categories><subj-group subj-group-type="heading"><subject>Article</subject></subj-group></article-categories><title>面向开放场景理解的遥感点云多模态融合与开放词汇技术综述</title><url>https://artdesignp.com/journal/TACS/3/5/10.61369/TACS.2026050048</url><author>王琼洁</author><pub-date pub-type="publication-year"><year>2026</year></pub-date><volume>3</volume><issue>5</issue><history><date date-type="pub"><published-time>2026-03-14</published-time></date></history><abstract>随着机载激光雷达、车载移动测量与无人机倾斜摄影等技术的快速发展，遥感场景中产生了海量高精度三维点云数据。传统遥感点云解译方法多建立在封闭类别监督范式之上，依赖大量点级标注，且在未知类别识别、跨场景迁移和复杂语义查询等方面存在明显局限。近年来，随着多模态预训练、开放词汇视觉理解以及大语言模型的发展，这些方式为遥感点云从&amp;ldquo;封闭集语义分割&amp;rdquo;走向&amp;ldquo;开放场景三维理解&amp;rdquo;提供了新的技术路径。基于此，本文围绕这一演进脉络，对相关研究进行系统梳理。首先，从遥感点云的大规模、稀疏性与弱纹理特征出发，概述了适用于大场景处理的点云表征骨干网络与自监督预训练范式。其次，本文重点总结了基于二维视觉- 语言模型的2D-to-3D 跨模态蒸馏与特征反投影方法与面向点云直接输入的三维语言模型构建方法这两类关键路线。最后，结合测绘遥感任务对几何精度、跨尺度一致性与工程部署的特殊要求，分析了现有方法在空间精确对齐、纯几何语义表达、长尾类别识别、数据集构建与可信落地等方面的主要瓶颈。</abstract><keywords>遥感点云,开放词汇分割,多模态融合,三维基础模型,3D-LLM,跨模态蒸馏</keywords></article-meta></front><body/><back><ref-list><ref id="B1" content-type="article"><label>1</label><element-citation publication-type="journal"><p>[1] 杨必胜, 董震. 点云智能研究进展与趋势[J]. 测绘学报,2019,48(12):1575-1585.[2] 胡伏原, 李晨露, 周涛, 等. 面向深度学习的三维点云补全算法综述[J]. 中国图象图形学报,2025,30(2):309-333. DOI:10.11834/jig.240124.[3] 潘洁晨, 邢帅, 曹家印, 等. 基于深度学习的航空点云语义分割研究进展[J]. 地球信息科学学报, 2025,27(9):1999-2020. DOI:10.12082/dqxxkx.2025.250151.[4] 张帅豪, 潘志刚. 遥感大模型: 综述与未来设想[J]. 遥感技术与应用,2025,40(1):1-13. DOI:10.11873/j.issn.1004-0323.2025.1.0001.[5]Hu Q, Yang B, Xie L, et al. Randla-net: Efficient semantic segmentation of large-scale point clouds[C]//Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2020: 11108-11117.[6]Choy C, Gwak J Y, Savarese S. 4d spatio-temporal convnets: Minkowski convolutional neural networks[C]//Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2019: 3075-3084.[7] 郑智鸿, 宋海川. 基于组对比学习的弱监督三维点云语义分割方法[J]. 华东师范大学学报（自然科学版）,2024(2):108-118. DOI:10.3969/j.issn.1000-5641.2024.02.012.[8]Yu X, Tang L, Rao Y, et al. Point-bert: Pre-training 3d point cloud transformers with masked point modeling[C]//Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2022: 19313-19322.[9]Pang Y, Tay E H F, Yuan L, et al. Masked autoencoders for 3d point cloud selfsupervisedlearning[J]. World Scientific Annual Review of Artificial Intelligence, 2023, 1: 2440001.[10]Peng S, Genova K, Jiang C, et al. Openscene: 3d scene understanding with open vocabularies[C]//Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2023: 815-824.[11]Radford A, Kim J W, Hallacy C, et al. Learning transferable visual models from natural language supervision[C]//International conference on machine learning. PmLR, 2021: 8748-8763.[12]Takmaz A, Fedele E, Sumner R W, et al. Openmask3d: Open-vocabulary 3d instance segmentation[J]. arXiv preprint arXiv:2306.13631, 2023.[13]guyen P, Ngo T D, Kalogerakis E, et al. Open3dis: Open-vocabulary 3d instance segmentation with 2d mask guidance[C]//Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2024: 4018-4028.[14] 付琨, 卢宛萱, 刘小煜, 等. 遥感基础模型发展综述与未来设想[J]. 遥感学报, 2024, 28(7).DOI:10.11834/jrs.20233313.[15]Hong Y, Zhen H, Chen P, et al. 3d-llm: Injecting the 3d world into large language models[J]. Advances in Neural Information Processing Systems, 2023, 36:20482-20494.[16]Xu R, Wang X, Wang T, et al. Pointllm: Empowering large language models to understand point clouds[C]//European Conference on Computer Vision. Cham: Springer Nature Switzerland, 2024: 131-147.</p><pub-id pub-id-type="doi"/></element-citation></ref></ref-list></back></article>
