<?xml version="1.1" encoding="utf-8"?>
<article xsi:noNamespaceSchemaLocation="http://jats.nlm.nih.gov/publishing/1.1/xsd/JATS-journalpublishing1-mathml3.xsd" dtd-version="1.1" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"><front><journal-meta><journal-id journal-id-type="publisher-id">TACS</journal-id><journal-title-group><journal-title>Technology and Application of Computer Science</journal-title></journal-title-group><issn>2998-8926</issn><eissn>2998-8934</eissn><publisher><publisher-name>Art and Technology</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.61369/TACS.2026050049</article-id><article-categories><subj-group subj-group-type="heading"><subject>Article</subject></subj-group></article-categories><title>聚焦视觉- 语言对齐：遥感多模态大模型技术综述与前沿展望</title><url>https://artdesignp.com/journal/TACS/3/5/10.61369/TACS.2026050049</url><author>王琼洁</author><pub-date pub-type="publication-year"><year>2026</year></pub-date><volume>3</volume><issue>5</issue><history><date date-type="pub"><published-time>2026-03-14</published-time></date></history><abstract>在视觉大模型快速演进的背景下，测绘遥感领域正由以分类、检测和分割为核心的单任务解译模式，逐步走向以视觉&amp;mdash; 语言对齐、指令跟随和自然语言交互为特征的多模态理解范式。相较于自然图像场景，遥感图像具有俯视视角、尺度跨度大、目标密集、跨传感器差异显著以及专业知识依赖强等特点，使通用视觉大模型难以直接迁移。围绕这一问题，本文从遥感视觉&amp;mdash; 文本对齐的基础与难点出发，系统梳理了遥感多模态大模型的数据构建、训练范式与核心架构演进。首先，概述了以CLIP 类模型为基础的遥感视觉&amp;mdash; 语言表征学习及其在领域持续预训练中的发展。其次，重点总结了遥感图文预训练数据、指令微调数据以及人机协同生成策略对模型能力形成的作用，并梳理了从全局图文理解到空间基础关联、再到多传感器统一建模的主要技术路线。最后，结合测绘遥感业务对空间定位精度、时空一致性、可靠性与工程可部署性的要求，分析了当前方法在遥感领域幻觉、细粒度定位、动态推理、评测体系和边缘部署等方面面临的主要问题。</abstract><keywords>遥感多模态大模型,视觉- 语言对齐,指令微调,空间基础关联,视觉问答,测绘遥感</keywords></article-meta></front><body/><back><ref-list><ref id="B1" content-type="article"><label>1</label><element-citation publication-type="journal"><p>[1] 龚健雅, 季顺平. 摄影测量与深度学习[J]. 测绘学报,2018,47(6):693-704.[2] 李德仁, 张良培, 夏桂松. 遥感大数据自动分析与数据挖掘[J]. 测绘学报,2014(12):1211-1216. DOI:10.13485/jc.nki1.1-20892.0140.187.[3] 燕琴, 顾海燕, 杨懿, 等. 智能遥感大模型研究进展与发展方向[J]. 测绘学报,2024,53(10):1967-1980. DOI:10.11947/j.AGCS.2024.20240053.3..[4] 张良培, 张乐飞, 袁强强. 遥感大模型: 进展与前瞻[J]. 武汉大学学报（信息科学版）,2023,48(10):1574-1581. DOI:10.13203/j.whugis20230341.[5] 张永军, 李彦胜, 党博, 等. 多模态遥感基础大模型: 研究现状与未来展望[J]. 测绘学报,2024,53(10):1942-1954. DOI:10.11947/j.AGCS.2024.20240019.[6]Radford A, Kim J W, Hallacy C, et al. Learning transferable visual models from natural language supervision[C]//International conference on machine learning. PmLR, 2021: 8748-8763.[7]Li J, Li D, Savarese S, et al. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models[C]//International conference on machine learning. PMLR, 2023: 19730-19742.[8]Liu F, Chen D, Guan Z, et al. Remoteclip: A vision language foundation model for remote sensing[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 1-16.[9]Zhang Z, Zhao T, Guo Y, et al. RS5M and GeoRSCLIP: A large-scale visionlanguage dataset and a large vision-language model for remote sensing[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 1-23.[10]Wang Z, Prabha R, Huang T, et al. SkyScript: A large and semantically diverse vision-language dataset for remote sensing[C]//Proceedings of the AAAI Conference on Artificial Intelligence. 2024, 38(6): 5805-5813.[11]Hu Y, Yuan J, Wen C, et al. Rsgpt: A remote sensing vision language model and benchmark[J]. ISPRS Journal of Photogrammetry and Remote Sensing, 2025, 224: 272-286.[12]Kuckreja K, Danish M S, Naseer M, et al. Geochat: Grounded large visionlanguage model for remote sensing[C]//Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2024: 27831-27840.[13]Zhang W, Cai M, Zhang T, et al. EarthGPT: A universal multimodal large language model for multisensor image comprehension in remote sensing domain[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 1-20.[14]Li X, Wen C, Hu Y, et al. Vision-language models in remote sensing: Current progress and future trends[J]. IEEE Geoscience and Remote Sensing Magazine, 2024, 12(2): 32-66.[15]Cha K, Yu D, Seo J. Pushing the limits of vision-language models in remote sensing without human annotations[J]. arXiv preprint arXiv:2409.07048, 2024.</p><pub-id pub-id-type="doi"/></element-citation></ref></ref-list></back></article>
