xarray 多维数据入门:从 combine='by_coords' 到维度与坐标

本文从一行常见代码出发,逐步拆解 xarray 中「维度」与「坐标」这两个核心概念,最后用一个鸡蛋盒的比喻帮你彻底搞懂它们。


目录


一、从一行代码说起

合并多个 NetCDF 文件时,你可能会写下这样的代码:

import xarray as xr

ds = xr.open_mfdataset("file_*.nc", combine='by_coords')
ds.to_netcdf("merged.nc")

看起来很简洁,但 combine='by_coords' 到底是什么意思?可以随便写吗?要回答这个问题,我们需要先理解 xarray 中两个最基础的概念——维度(Dimension) 和 坐标(Coordinate)。


二、combine='by_coords' 到底在做什么

2.1 combine 参数的合法值

combine 是一个固定选项参数,告诉 xarray 如何把多个文件合并成一个 Dataset。它不能随便写,只有以下合法值:

值含义底层调用
'by_coords'根据坐标值自动排序、对齐后合并xr.combine_by_coords()
'nested'按文件顺序(嵌套列表结构)拼接,需配合 concat_dimxr.combine_nested()
'compat'⚠️ 已弃用(deprecated),旧版默认行为—

写成 combine='abc' 会直接抛出 ValueError。

2.2 'by_coords' 的核心逻辑

读取每个文件的坐标(coords),按坐标值排序,然后沿不同维度自动 concat 或 merge。

关键要求:每个文件必须有能区分彼此的坐标变量(通常是 time、lat、lon 等)。

2.3 典型场景

✅ 场景 1:按时间切分(最常见)
file_2024.nc  →  time: 2024-01-01 ~ 2024-12-31
file_2025.nc  →  time: 2025-01-01 ~ 2025-12-31
file_2026.nc  →  time: 2026-01-01 ~ 2026-06-30
ds = xr.open_mfdataset("file_*.nc", combine='by_coords')
# xarray 读取每个文件的 time 坐标 → 排序 → 沿 time 维度 concat
# 结果: time 从 2024-01-01 连续到 2026-06-30
✅ 场景 2:按空间区域切分
file_north.nc  →  lat: 30°N ~ 90°N
file_south.nc  →  lat: -90°S ~ 30°N
ds = xr.open_mfdataset("file_*.nc", combine='by_coords')
# xarray 发现 lat 坐标不同 → 沿 lat 维度 concat
# 结果: lat 从 -90 到 90 完整覆盖
✅ 场景 3:多维网格切分(时间 × 空间)
file_2024_north.nc  →  time: 2024, lat: 30~90
file_2024_south.nc  →  time: 2024, lat: -90~30
file_2025_north.nc  →  time: 2025, lat: 30~90
file_2025_south.nc  →  time: 2025, lat: -90~30
ds = xr.open_mfdataset("file_*.nc", combine='by_coords')
# xarray 自动识别:
#   - time 不同 → 沿 time concat
#   - lat 不同  → 沿 lat concat
# 最终得到完整的 (time × lat × lon) 数据集

这就是 by_coords 最强大的地方——自动推断多维拼接方式。

❌ 场景 4:文件没有区分性坐标(会出问题)
file_a.nc  →  dims: (x: 100, y: 100)  # 没有坐标,只有维度名
file_b.nc  →  dims: (x: 100, y: 100)  # 完全一样
ds = xr.open_mfdataset("file_*.nc", combine='by_coords')
# ⚠️ 报错或结果不对!xarray 无法通过坐标区分这两个文件该怎么拼

此时应改用 'nested':

ds = xr.open_mfdataset("file_*.nc", combine='nested', concat_dim='x')
# 明确告诉 xarray:按文件顺序,沿 x 维度拼接

2.4 'by_coords' vs 'nested' 对比

# ✅ by_coords:自动按坐标排序合并(不需要指定 concat_dim)
ds = xr.open_mfdataset("file_*.nc", combine='by_coords')

# ✅ nested:手动指定拼接维度(按文件列表顺序)
ds = xr.open_mfdataset("file_*.nc", combine='nested', concat_dim='time')

# ✅ 多维 nested:文件组织成嵌套列表
ds = xr.open_mfdataset(
    [["file_2024_north.nc", "file_2024_south.nc"],
     ["file_2025_north.nc", "file_2025_south.nc"]],
    combine='nested',
    concat_dim=['time', 'lat']
)

2.5 小结

问题回答
by_coords 由什么决定?由文件内部的坐标变量(如 time, lat, lon)决定,xarray 自动读取并排序
可以随便写吗?❌ 不可以。只能写 'by_coords'、'nested'、'compat'(已弃用)
什么时候用 by_coords?文件有明确的、不重叠的坐标(如不同时间段、不同空间区域)
什么时候用 nested?文件没有区分性坐标,或你想手动控制拼接顺序和维度

💡 经验法则:如果你的 .nc 文件里有 time 坐标且每个文件对应不同时间段,直接用默认的 combine='by_coords' 就好。


三、维度与坐标:联系与区别

理解了 by_coords 依赖坐标工作之后,我们来正式认识这两个概念。

3.1 一句话定义

概念本质类比
维度 (Dimension)一个名字 + 长度,描述数据"有几个轴、每个轴多长"表格的行号/列号(0, 1, 2, …)
坐标 (Coordinate)附着在维度上的具体标签值,描述"每个位置是什么"表格的日期列、城市名(2024-01-01, 北京, …)

维度回答 “多大”,坐标回答 “是什么”。

3.2 有坐标 vs 无坐标

import numpy as np
import xarray as xr

data = np.random.rand(3, 4)  # 3行 × 4列

# ===== ① 只有维度,没有坐标 =====
da1 = xr.DataArray(data, dims=['x', 'y'])
print(da1)
<xarray.DataArray (x: 3, y: 4)>
array([[0.54, 0.71, 0.60, 0.54],
       [0.42, 0.65, 0.44, 0.89],
       [0.96, 0.38, 0.79, 0.53]])
Dimensions without coordinates:   ← ⚠️ 注意这行
    x: 3
    y: 4

只知道 x 轴有 3 个位置、y 轴有 4 个位置,但不知道每个位置代表什么。

# ===== ② 有维度 + 有坐标 =====
da2 = xr.DataArray(
    data,
    dims=['time', 'city'],
    coords={
        'time': ['2024-01', '2024-02', '2024-03'],
        'city': ['北京', '上海', '广州', '深圳'],
    }
)
print(da2)
<xarray.DataArray (time: 3, city: 4)>
array([[0.54, 0.71, 0.60, 0.54],
       [0.42, 0.65, 0.44, 0.89],
       [0.96, 0.38, 0.79, 0.53]])
Coordinates:                      ← ✅ 有了!
  * time   (time) <U7 '2024-01' '2024-02' '2024-03'
  * city   (city) <U2 '北京' '上海' '广州' '深圳'
Dimensions without coordinates:
    *empty*

3.3 核心区别

① 索引方式不同
# 无坐标 → 只能用位置索引 .isel (integer select)
da1.isel(x=0, y=2)       # ✅ 第0行第2列
da1.sel(x='2024-01')     # ❌ 报错!没有坐标,无法按标签选

# 有坐标 → 既可以用位置,也可以用标签
da2.isel(time=0)         # ✅ 位置索引
da2.sel(time='2024-01')  # ✅ 标签索引 ← 坐标的核心价值
da2.sel(city='上海')      # ✅
② 对齐 (Alignment) 能力不同
a = xr.DataArray([10, 20, 30], dims='x', coords={'x': [1, 2, 3]})
b = xr.DataArray([40, 50, 60], dims='x', coords={'x': [2, 3, 4]})

a + b
# 自动按坐标对齐:
# x=1: 10 + NaN = NaN
# x=2: 20 + 40  = 60
# x=3: 30 + 50  = 80
# x=4: NaN + 60 = NaN

没有坐标的话,xarray 只能按位置硬拼,无法做语义对齐。

③ 合并文件时的作用不同
# combine='by_coords' 靠的就是坐标来排序和拼接
# 如果文件只有维度没有坐标 → by_coords 无法工作

四、坐标的两种类型

这是很多人忽略的重点:

ds = xr.Dataset(
    data_vars={
        'temperature': (['time', 'y', 'x'], np.random.rand(3, 4, 5)),
    },
    coords={
        # ---- 维度坐标 (Dimension Coordinate) ----
        # 1D,名字与维度同名,带 * 号
        'time': pd.date_range('2024-01-01', periods=3),
        'y':     np.arange(4),
        'x':     np.arange(5),

        # ---- 非维度坐标 (Non-dimension Coordinate) ----
        # 可以是多维的,名字与维度不同名
        'lat':  (['y', 'x'], np.random.rand(4, 5)),   # 二维经纬度
        'lon':  (['y', 'x'], np.random.rand(4, 5)),
    }
)
print(ds)
<xarray.Dataset>
Dimensions:      (time: 3, y: 4, x: 5)
Coordinates:
  * time         (time) datetime64[ns] 2024-01-01 2024-01-02 2024-01-03
  * y            (y) int64 0 1 2 3
  * x            (x) int64 0 1 2 3 4
    lat          (y, x) float64 ...          ← 非维度坐标(2D)
    lon          (y, x) float64 ...          ← 非维度坐标(2D)
Data variables:
    temperature  (time, y, x) float64 ...
类型特征作用
维度坐标1D,名字 = 维度名,打印时带 *支持 .sel() 标签索引、自动对齐、by_coords 合并
非维度坐标可多维,名字 ≠ 维度名提供辅助描述信息(如曲线网格的经纬度),不能直接用于 .sel()

一个完整的现实例子

# 气象数据:3天 × 4个站点
ds = xr.Dataset(
    {
        'temp':     (['time', 'station'], np.random.rand(3, 4) * 30),
        'humidity': (['time', 'station'], np.random.rand(3, 4) * 100),
    },
    coords={
        'time':    pd.date_range('2024-07-01', periods=3),          # 维度坐标
        'station': ['北京', '上海', '广州', '深圳'],                  # 维度坐标
        'lat':     ('station', [39.9, 31.2, 23.1, 22.5]),          # 非维度坐标
        'lon':     ('station', [116.4, 121.5, 113.3, 114.1]),      # 非维度坐标
    }
)
维度 (Dimensions):
    time: 3        ← 只是"有3个时间步"
    station: 4     ← 只是"有4个站点"

坐标 (Coordinates):
  * time    → 2024-07-01, 2024-07-02, 2024-07-03   ← 具体是哪3天
  * station → 北京, 上海, 广州, 深圳                  ← 具体是哪4站
    lat     → 39.9, 31.2, 23.1, 22.5               ← 附加地理信息
    lon     → 116.4, 121.5, 113.3, 114.1
ds.sel(time='2024-07-02', station='上海')   # ✅ 靠坐标
ds.isel(time=1, station=1)                  # ✅ 靠维度位置
ds.sel(lat=31.2)                            # ❌ 非维度坐标不能直接 sel

五、给 10 岁小朋友的解释:鸡蛋盒比喻 🥚

如果上面的内容还是有点抽象,我们换一个方式。

想象一个鸡蛋盒

你打开冰箱,拿出一个鸡蛋盒:

        第1格   第2格   第3格   第4格   第5格   第6格
第1排  [ 🥚 ]  [ 🥚 ]  [ 🥚 ]  [ 🥚 ]  [ 🥚 ]  [ 🥚 ]
第2排  [ 🥚 ]  [ 🥚 ]  [ 🥚 ]  [ 🥚 ]  [ 🥚 ]  [ 🥚 ]

现在我问你两个问题:

“这个盒子有多大?”
→ 2排 × 6格。这就是维度。

“第1排第3格里放的是什么?”
→ 呃……一个鸡蛋?哪个鸡蛋?周一买的还是周二买的?不知道!

光知道"多大",不知道"是什么",是不是很没用?

给鸡蛋盒贴上标签 🏷️

妈妈在盒子上贴了标签:

         周一    周二    周三    周四    周五    周六
土鸡蛋  [ 🥚 ]  [ 🥚 ]  [ 🥚 ]  [ 🥚 ]  [ 🥚 ]  [ 🥚 ]
洋鸡蛋  [ 🥚 ]  [ 🥚 ]  [ 🥚 ]  [ 🥚 ]  [ 🥚 ]  [ 🥚 ]

现在再问:

“周三的土鸡蛋是哪个?”
→ 你一眼就能指出来!✅

这些 "周一、周二……“和"土鸡蛋、洋鸡蛋”,就是坐标。

对照表

鸡蛋盒意思
维度2排 × 6格盒子有多大(骨架)
坐标周一~周六、土鸡蛋/洋鸡蛋每个格子是什么(标签)

🦴 维度是骨架:没有它,盒子都不存在,鸡蛋没地方放。
🏷️ 坐标是标签:没有它,盒子还在,但你分不清哪个是哪个。

换成气象数据 🌦️

科学家记录天气,就像填一个超级大表格:

              北京    上海    广州    深圳
7月1日       35°C    33°C    34°C    32°C
7月2日       36°C    31°C    33°C    33°C
7月3日       34°C    32°C    35°C    31°C
  • 维度:3天 × 4个城市 → 告诉你表格有多大
  • 坐标:「7月1日、7月2日、7月3日」和「北京、上海、广州、深圳」→ 告诉你每行每列是什么

没有坐标会怎样?😱

如果科学家忘了写标签,表格就变成了:

              ???     ???     ???     ???
???          35      33      34      32
???          36      31      33      33
???          34      32      35      31

你看着这堆数字,完全不知道:

“35度是哪里?哪天?” 🤷

数据还在,但变成了天书。

一句话记住 ✨

维度告诉你:这个盒子有 几排几格。
坐标告诉你:这一格放的是 周三的土鸡蛋。

没有盒子,鸡蛋没地方放;
没有标签,你分不清哪颗是哪颗。


六、总结

结构全景图

┌─────────────────────────────────────────────────┐
│              DataArray / Dataset                 │
│                                                  │
│   维度 (Dimension)                               │
│   ┌──────────────────────────────┐               │
│   │  time: 3    station: 4      │  ← 名字+长度  │
│   └──────────────────────────────┘               │
│          │                │                      │
│          ▼                ▼                      │
│   坐标 (Coordinate)                              │
│   ┌──────────────────────────────┐               │
│   │ *time:    07-01,07-02,07-03 │  ← 维度坐标   │
│   │ *station: 北京,上海,广州,深圳 │  ← 维度坐标   │
│   │  lat:     39.9, 31.2, ...   │  ← 非维度坐标 │
│   │  lon:     116.4, 121.5, ... │  ← 非维度坐标 │
│   └──────────────────────────────┘               │
│                                                  │
│   数据变量 (Data Variable)                        │
│   ┌──────────────────────────────┐               │
│   │  temp:     (time, station)   │  ← 实际数值   │
│   │  humidity: (time, station)   │               │
│   └──────────────────────────────┘               │
└─────────────────────────────────────────────────┘

核心要点回顾

要点说明
维度是骨架定义数据的形状和大小,没有维度数据无法存在
坐标是标签赋予每个位置具体含义,没有坐标数据失去语义
by_coords 依赖坐标它通过读取坐标值来自动排序和拼接文件
combine 不能随便写只有 'by_coords'、'nested'、'compat'(已弃用)
维度坐标 vs 非维度坐标前者可 .sel()、可对齐;后者仅提供辅助信息

最终一句话:维度是骨架(多大),坐标是标签(是什么)。没有骨架,数据无处安放;没有标签,数据形同天书。而 combine='by_coords',正是 xarray 利用这些"标签"来自动拼合多文件数据的魔法。

更多推荐