NumPy & pandas 复习
讲义:
- pandas_1/2/3 讲义:
- Python + NumPy 教程(Data 100 推荐):https://cs231n.github.io/python-numpy-tutorial/
1. NumPy 快速复习
1.1 ndarray 创建与属性
import numpy as np
arr = np.array([[1, 2, 3], [4, 5, 6]])
arr.shape # (2, 3)
arr.dtype # dtype('int64')
arr.size # 6
np.zeros((2, 3)) # 全 0 矩阵
np.ones(4) # 全 1 向量
np.arange(0, 10, 2) # array([0, 2, 4, 6, 8])
np.linspace(0, 1, 5) # array([0. , 0.25, 0.5 , 0.75, 1. ])
np.random.randn(3, 3) # 标准正态随机矩阵1.2 Indexing 与 Slicing(索引与切片)
arr[0, 1] # 第 0 行第 1 列(0-based)
arr[0, :] # 第 0 行 → array([1, 2, 3])
arr[:, 1] # 第 1 列 → array([2, 5])
arr[1:, :2] # 子矩阵:第 1 行起、前 2 列- Slicing 不含右端点:
arr[1:3]取第 1、2 行 - 与 R 不同,Python/NumPy 索引从 0 开始
1.3 广播 Broadcasting
a = np.array([[1, 2, 3], [4, 5, 6]])
a + 10 # 标量广播
a * np.array([1, 10, 100]) # (2,3) * (3,) → 按行逐列乘
a.T + np.array([1, 2, 3]) # (3,2) + (3,) → 按列逐行加- Broadcasting 规则:从尾部维度对齐,维度相等或为 1 时可扩展
- 无法对齐时报错,例如
(2,3)与(2,)相乘
1.4 Vectorization(向量化)
# 优先 Vectorization,避免 Python 循环逐元素操作
np.where(a > 3, a, 0) # 条件选择
np.sum(a, axis=0) # 按列求和 → array([5, 7, 9])
np.mean(a), np.std(a)- 矩阵乘法:
np.dot(a, b)或a @ b - pandas 底层就是 NumPy:
df.to_numpy()取底层 ndarray
2. pandas 复习(Data 100 pandas_1/2/3)
2.1 DataFrame 构建与 Index(pandas_1)
import pandas as pd
# 从 dict 构建
df = pd.DataFrame({
"name": ["Alice", "Bob", "Charlie"],
"score": [88, 92, 75],
})
# 从 list + columns 构建
pd.DataFrame([[1, 2], [3, 4]], columns=["a", "b"])
# 自定义 index
df.index = ["u1", "u2", "u3"]属性:df.index、df.columns、df.shape
提取数据(Data 100 的三种 Indexing 方式):
df.head(2) / df.tail(1)
df.loc["u1"] # 标签索引,含右端点
df.iloc[0] # 整数位置索引,不含右端点
df["score"] # 单列 → Series
df[["name", "score"]] # 多列 → DataFrame
df[:2] # 行切片(整数位置)2.2 Conditional Selection 与列操作(pandas_2)
df[df["score"] >= 80] # 条件选择
df[(df["score"] >= 80) & (df["score"] < 90)] # 多条件必须加括号
df[df["name"].str.startswith("A")]增删改列:
df["grade"] = ["A", "A", "B"] # 新增列
df["score2"] = df["score"] * 2 # 由已有列生成
df = df.drop(columns=["score2"]) # 删列(返回新 df)
df.loc["u1", "score"] = 90 # 修改单元格常用工具函数:
df.shape / df.size
df.describe() # 描述统计
df.sample(2) # 随机抽样
df["score"].value_counts() # 频数统计
df["score"].unique() # 去重
df.sort_values("score", ascending=False)自定义排序(Data 100 三种做法):
# 方式 1:临时列
df["name_len"] = df["name"].str.len()
df.sort_values("name_len").drop(columns="name_len")
# 方式 2:key 参数
df.sort_values("name", key=lambda s: s.str.len())
# 方式 3:map 建立排序键
grade_rank = {"A": 0, "B": 1, "C": 2}
df.sort_values("grade", key=lambda s: s.map(grade_rank))2.3 Groupby、Pivot Table 与 Merge(pandas_3)
df.groupby("grade")["score"].mean()
df.groupby("grade").agg({"score": ["mean", "max"]})
df.groupby("grade").filter(lambda g: len(g) >= 2)
df.groupby("grade")["score"].agg(lambda x: x.max() - x.min())
df.pivot_table(index="grade", values="score", aggfunc="mean")
pd.merge(df, df2, on="id", how="left") # inner / left / right / outergroupby后非数值列默认不参与数值聚合(nuisance columns),可先显式选列,或用numeric_only=Truepivot_table的aggfunc默认"mean"
3. 坑
[]索引上下文依赖:传单个列名 → Series;传列表 → DataFrame;传切片 → 按行- 多条件布尔选择必须用
&/|且每个条件加括号,不能用and/or .loc含右端点,.iloc不含;R 转 Python 时尤其容易混- 链式赋值(chained assignment)可能触发
SettingWithCopyWarning:先.copy(),或直接用.loc groupby聚合非数值列报错(nuisance columns),显式选列或加numeric_only=True- Broadcasting 维度对不齐时静默出错概率高,先
print(a.shape)再操作
4. 函数速查
| 函数 / 方法 | 作用 |
|---|---|
np.arange / np.linspace | 生成等距序列 |
arr.reshape / arr.T | 变形 / 转置 |
np.dot / @ | 矩阵乘法 |
np.where | 条件向量化选择 |
df.loc / df.iloc | 标签 / 整数位置索引 |
df.groupby().agg() | 分组聚合 |
df.pivot_table | 透视表 |
pd.merge | 表连接 |
df.sort_values | 排序(可用 key= 自定义) |
df.describe / sample / value_counts | 描述统计 / 抽样 / 频数 |