xarray 多维数据入门:从 `combine=‘by_coords‘` 到维度与坐标
xarray 多维数据入门:从 combine='by_coords' 到维度与坐标
本文从一行常见代码出发,逐步拆解 xarray 中「维度」与「坐标」这两个核心概念,最后用一个鸡蛋盒的比喻帮你彻底搞懂它们。
目录
一、从一行代码说起
合并多个 NetCDF 文件时,你可能会写下这样的代码:
import xarray as xr
ds = xr.open_mfdataset("file_*.nc", combine='by_coords')
ds.to_netcdf("merged.nc")
看起来很简洁,但 combine='by_coords' 到底是什么意思?可以随便写吗?要回答这个问题,我们需要先理解 xarray 中两个最基础的概念——维度(Dimension) 和 坐标(Coordinate)。
二、combine='by_coords' 到底在做什么
2.1 combine 参数的合法值
combine 是一个固定选项参数,告诉 xarray 如何把多个文件合并成一个 Dataset。它不能随便写,只有以下合法值:
| 值 | 含义 | 底层调用 |
|---|---|---|
'by_coords' | 根据坐标值自动排序、对齐后合并 | xr.combine_by_coords() |
'nested' | 按文件顺序(嵌套列表结构)拼接,需配合 concat_dim | xr.combine_nested() |
'compat' | ⚠️ 已弃用(deprecated),旧版默认行为 | — |
写成
combine='abc'会直接抛出ValueError。
2.2 'by_coords' 的核心逻辑
读取每个文件的坐标(coords),按坐标值排序,然后沿不同维度自动 concat 或 merge。
关键要求:每个文件必须有能区分彼此的坐标变量(通常是 time、lat、lon 等)。
2.3 典型场景
✅ 场景 1:按时间切分(最常见)
file_2024.nc → time: 2024-01-01 ~ 2024-12-31
file_2025.nc → time: 2025-01-01 ~ 2025-12-31
file_2026.nc → time: 2026-01-01 ~ 2026-06-30
ds = xr.open_mfdataset("file_*.nc", combine='by_coords')
# xarray 读取每个文件的 time 坐标 → 排序 → 沿 time 维度 concat
# 结果: time 从 2024-01-01 连续到 2026-06-30
✅ 场景 2:按空间区域切分
file_north.nc → lat: 30°N ~ 90°N
file_south.nc → lat: -90°S ~ 30°N
ds = xr.open_mfdataset("file_*.nc", combine='by_coords')
# xarray 发现 lat 坐标不同 → 沿 lat 维度 concat
# 结果: lat 从 -90 到 90 完整覆盖
✅ 场景 3:多维网格切分(时间 × 空间)
file_2024_north.nc → time: 2024, lat: 30~90
file_2024_south.nc → time: 2024, lat: -90~30
file_2025_north.nc → time: 2025, lat: 30~90
file_2025_south.nc → time: 2025, lat: -90~30
ds = xr.open_mfdataset("file_*.nc", combine='by_coords')
# xarray 自动识别:
# - time 不同 → 沿 time concat
# - lat 不同 → 沿 lat concat
# 最终得到完整的 (time × lat × lon) 数据集
这就是 by_coords 最强大的地方——自动推断多维拼接方式。
❌ 场景 4:文件没有区分性坐标(会出问题)
file_a.nc → dims: (x: 100, y: 100) # 没有坐标,只有维度名
file_b.nc → dims: (x: 100, y: 100) # 完全一样
ds = xr.open_mfdataset("file_*.nc", combine='by_coords')
# ⚠️ 报错或结果不对!xarray 无法通过坐标区分这两个文件该怎么拼
此时应改用 'nested':
ds = xr.open_mfdataset("file_*.nc", combine='nested', concat_dim='x')
# 明确告诉 xarray:按文件顺序,沿 x 维度拼接
2.4 'by_coords' vs 'nested' 对比
# ✅ by_coords:自动按坐标排序合并(不需要指定 concat_dim)
ds = xr.open_mfdataset("file_*.nc", combine='by_coords')
# ✅ nested:手动指定拼接维度(按文件列表顺序)
ds = xr.open_mfdataset("file_*.nc", combine='nested', concat_dim='time')
# ✅ 多维 nested:文件组织成嵌套列表
ds = xr.open_mfdataset(
[["file_2024_north.nc", "file_2024_south.nc"],
["file_2025_north.nc", "file_2025_south.nc"]],
combine='nested',
concat_dim=['time', 'lat']
)
2.5 小结
| 问题 | 回答 |
|---|---|
by_coords 由什么决定? | 由文件内部的坐标变量(如 time, lat, lon)决定,xarray 自动读取并排序 |
| 可以随便写吗? | ❌ 不可以。只能写 'by_coords'、'nested'、'compat'(已弃用) |
什么时候用 by_coords? | 文件有明确的、不重叠的坐标(如不同时间段、不同空间区域) |
什么时候用 nested? | 文件没有区分性坐标,或你想手动控制拼接顺序和维度 |
💡 经验法则:如果你的
.nc文件里有time坐标且每个文件对应不同时间段,直接用默认的combine='by_coords'就好。
三、维度与坐标:联系与区别
理解了 by_coords 依赖坐标工作之后,我们来正式认识这两个概念。
3.1 一句话定义
| 概念 | 本质 | 类比 |
|---|---|---|
| 维度 (Dimension) | 一个名字 + 长度,描述数据"有几个轴、每个轴多长" | 表格的行号/列号(0, 1, 2, …) |
| 坐标 (Coordinate) | 附着在维度上的具体标签值,描述"每个位置是什么" | 表格的日期列、城市名(2024-01-01, 北京, …) |
维度回答 “多大”,坐标回答 “是什么”。
3.2 有坐标 vs 无坐标
import numpy as np
import xarray as xr
data = np.random.rand(3, 4) # 3行 × 4列
# ===== ① 只有维度,没有坐标 =====
da1 = xr.DataArray(data, dims=['x', 'y'])
print(da1)
<xarray.DataArray (x: 3, y: 4)>
array([[0.54, 0.71, 0.60, 0.54],
[0.42, 0.65, 0.44, 0.89],
[0.96, 0.38, 0.79, 0.53]])
Dimensions without coordinates: ← ⚠️ 注意这行
x: 3
y: 4
只知道
x轴有 3 个位置、y轴有 4 个位置,但不知道每个位置代表什么。
# ===== ② 有维度 + 有坐标 =====
da2 = xr.DataArray(
data,
dims=['time', 'city'],
coords={
'time': ['2024-01', '2024-02', '2024-03'],
'city': ['北京', '上海', '广州', '深圳'],
}
)
print(da2)
<xarray.DataArray (time: 3, city: 4)>
array([[0.54, 0.71, 0.60, 0.54],
[0.42, 0.65, 0.44, 0.89],
[0.96, 0.38, 0.79, 0.53]])
Coordinates: ← ✅ 有了!
* time (time) <U7 '2024-01' '2024-02' '2024-03'
* city (city) <U2 '北京' '上海' '广州' '深圳'
Dimensions without coordinates:
*empty*
3.3 核心区别
① 索引方式不同
# 无坐标 → 只能用位置索引 .isel (integer select)
da1.isel(x=0, y=2) # ✅ 第0行第2列
da1.sel(x='2024-01') # ❌ 报错!没有坐标,无法按标签选
# 有坐标 → 既可以用位置,也可以用标签
da2.isel(time=0) # ✅ 位置索引
da2.sel(time='2024-01') # ✅ 标签索引 ← 坐标的核心价值
da2.sel(city='上海') # ✅
② 对齐 (Alignment) 能力不同
a = xr.DataArray([10, 20, 30], dims='x', coords={'x': [1, 2, 3]})
b = xr.DataArray([40, 50, 60], dims='x', coords={'x': [2, 3, 4]})
a + b
# 自动按坐标对齐:
# x=1: 10 + NaN = NaN
# x=2: 20 + 40 = 60
# x=3: 30 + 50 = 80
# x=4: NaN + 60 = NaN
没有坐标的话,xarray 只能按位置硬拼,无法做语义对齐。
③ 合并文件时的作用不同
# combine='by_coords' 靠的就是坐标来排序和拼接
# 如果文件只有维度没有坐标 → by_coords 无法工作
四、坐标的两种类型
这是很多人忽略的重点:
ds = xr.Dataset(
data_vars={
'temperature': (['time', 'y', 'x'], np.random.rand(3, 4, 5)),
},
coords={
# ---- 维度坐标 (Dimension Coordinate) ----
# 1D,名字与维度同名,带 * 号
'time': pd.date_range('2024-01-01', periods=3),
'y': np.arange(4),
'x': np.arange(5),
# ---- 非维度坐标 (Non-dimension Coordinate) ----
# 可以是多维的,名字与维度不同名
'lat': (['y', 'x'], np.random.rand(4, 5)), # 二维经纬度
'lon': (['y', 'x'], np.random.rand(4, 5)),
}
)
print(ds)
<xarray.Dataset>
Dimensions: (time: 3, y: 4, x: 5)
Coordinates:
* time (time) datetime64[ns] 2024-01-01 2024-01-02 2024-01-03
* y (y) int64 0 1 2 3
* x (x) int64 0 1 2 3 4
lat (y, x) float64 ... ← 非维度坐标(2D)
lon (y, x) float64 ... ← 非维度坐标(2D)
Data variables:
temperature (time, y, x) float64 ...
| 类型 | 特征 | 作用 |
|---|---|---|
| 维度坐标 | 1D,名字 = 维度名,打印时带 * | 支持 .sel() 标签索引、自动对齐、by_coords 合并 |
| 非维度坐标 | 可多维,名字 ≠ 维度名 | 提供辅助描述信息(如曲线网格的经纬度),不能直接用于 .sel() |
一个完整的现实例子
# 气象数据:3天 × 4个站点
ds = xr.Dataset(
{
'temp': (['time', 'station'], np.random.rand(3, 4) * 30),
'humidity': (['time', 'station'], np.random.rand(3, 4) * 100),
},
coords={
'time': pd.date_range('2024-07-01', periods=3), # 维度坐标
'station': ['北京', '上海', '广州', '深圳'], # 维度坐标
'lat': ('station', [39.9, 31.2, 23.1, 22.5]), # 非维度坐标
'lon': ('station', [116.4, 121.5, 113.3, 114.1]), # 非维度坐标
}
)
维度 (Dimensions):
time: 3 ← 只是"有3个时间步"
station: 4 ← 只是"有4个站点"
坐标 (Coordinates):
* time → 2024-07-01, 2024-07-02, 2024-07-03 ← 具体是哪3天
* station → 北京, 上海, 广州, 深圳 ← 具体是哪4站
lat → 39.9, 31.2, 23.1, 22.5 ← 附加地理信息
lon → 116.4, 121.5, 113.3, 114.1
ds.sel(time='2024-07-02', station='上海') # ✅ 靠坐标
ds.isel(time=1, station=1) # ✅ 靠维度位置
ds.sel(lat=31.2) # ❌ 非维度坐标不能直接 sel
五、给 10 岁小朋友的解释:鸡蛋盒比喻 🥚
如果上面的内容还是有点抽象,我们换一个方式。
想象一个鸡蛋盒
你打开冰箱,拿出一个鸡蛋盒:
第1格 第2格 第3格 第4格 第5格 第6格
第1排 [ 🥚 ] [ 🥚 ] [ 🥚 ] [ 🥚 ] [ 🥚 ] [ 🥚 ]
第2排 [ 🥚 ] [ 🥚 ] [ 🥚 ] [ 🥚 ] [ 🥚 ] [ 🥚 ]
现在我问你两个问题:
“这个盒子有多大?”
→ 2排 × 6格。这就是维度。
“第1排第3格里放的是什么?”
→ 呃……一个鸡蛋?哪个鸡蛋?周一买的还是周二买的?不知道!
光知道"多大",不知道"是什么",是不是很没用?
给鸡蛋盒贴上标签 🏷️
妈妈在盒子上贴了标签:
周一 周二 周三 周四 周五 周六
土鸡蛋 [ 🥚 ] [ 🥚 ] [ 🥚 ] [ 🥚 ] [ 🥚 ] [ 🥚 ]
洋鸡蛋 [ 🥚 ] [ 🥚 ] [ 🥚 ] [ 🥚 ] [ 🥚 ] [ 🥚 ]
现在再问:
“周三的土鸡蛋是哪个?”
→ 你一眼就能指出来!✅
这些 "周一、周二……“和"土鸡蛋、洋鸡蛋”,就是坐标。
对照表
| 鸡蛋盒 | 意思 | |
|---|---|---|
| 维度 | 2排 × 6格 | 盒子有多大(骨架) |
| 坐标 | 周一~周六、土鸡蛋/洋鸡蛋 | 每个格子是什么(标签) |
🦴 维度是骨架:没有它,盒子都不存在,鸡蛋没地方放。
🏷️ 坐标是标签:没有它,盒子还在,但你分不清哪个是哪个。
换成气象数据 🌦️
科学家记录天气,就像填一个超级大表格:
北京 上海 广州 深圳
7月1日 35°C 33°C 34°C 32°C
7月2日 36°C 31°C 33°C 33°C
7月3日 34°C 32°C 35°C 31°C
- 维度:3天 × 4个城市 → 告诉你表格有多大
- 坐标:「7月1日、7月2日、7月3日」和「北京、上海、广州、深圳」→ 告诉你每行每列是什么
没有坐标会怎样?😱
如果科学家忘了写标签,表格就变成了:
??? ??? ??? ???
??? 35 33 34 32
??? 36 31 33 33
??? 34 32 35 31
你看着这堆数字,完全不知道:
“35度是哪里?哪天?” 🤷
数据还在,但变成了天书。
一句话记住 ✨
维度告诉你:这个盒子有 几排几格。
坐标告诉你:这一格放的是 周三的土鸡蛋。没有盒子,鸡蛋没地方放;
没有标签,你分不清哪颗是哪颗。
六、总结
结构全景图
┌─────────────────────────────────────────────────┐
│ DataArray / Dataset │
│ │
│ 维度 (Dimension) │
│ ┌──────────────────────────────┐ │
│ │ time: 3 station: 4 │ ← 名字+长度 │
│ └──────────────────────────────┘ │
│ │ │ │
│ ▼ ▼ │
│ 坐标 (Coordinate) │
│ ┌──────────────────────────────┐ │
│ │ *time: 07-01,07-02,07-03 │ ← 维度坐标 │
│ │ *station: 北京,上海,广州,深圳 │ ← 维度坐标 │
│ │ lat: 39.9, 31.2, ... │ ← 非维度坐标 │
│ │ lon: 116.4, 121.5, ... │ ← 非维度坐标 │
│ └──────────────────────────────┘ │
│ │
│ 数据变量 (Data Variable) │
│ ┌──────────────────────────────┐ │
│ │ temp: (time, station) │ ← 实际数值 │
│ │ humidity: (time, station) │ │
│ └──────────────────────────────┘ │
└─────────────────────────────────────────────────┘
核心要点回顾
| 要点 | 说明 |
|---|---|
| 维度是骨架 | 定义数据的形状和大小,没有维度数据无法存在 |
| 坐标是标签 | 赋予每个位置具体含义,没有坐标数据失去语义 |
by_coords 依赖坐标 | 它通过读取坐标值来自动排序和拼接文件 |
combine 不能随便写 | 只有 'by_coords'、'nested'、'compat'(已弃用) |
| 维度坐标 vs 非维度坐标 | 前者可 .sel()、可对齐;后者仅提供辅助信息 |
最终一句话:维度是骨架(多大),坐标是标签(是什么)。没有骨架,数据无处安放;没有标签,数据形同天书。而
combine='by_coords',正是 xarray 利用这些"标签"来自动拼合多文件数据的魔法。
更多推荐


所有评论(0)