2026-08-05-Note · Data 100: Visualization I(分布可视化)

课程:UC Berkeley Data 100(https://ds100.org/)

讲义:Visualization I - Data 100: Principles and Techniques of Data Science

0. 环境准备与数据读取

Cell 2:导入库并读取世界银行数据

import pandas as pd
import numpy as np
import warnings 
 
warnings.filterwarnings("ignore", "use_inf_as_na") # Supresses distracting deprecation warnings
 
wb = pd.read_csv("data/world_bank.csv", index_col=0)
wb.head()
  • import warnings:导入 warnings 模块,用来屏蔽干扰性警告;
  • warnings.filterwarnings("ignore", "use_inf_as_na"):忽略关于 use_inf_as_na 的弃用(deprecation)警告——pandas 未来版本会把 inf 当缺失值处理,这里先静音;
  • wb = pd.read_csv("data/world_bank.csv", index_col=0):读取 data/ 目录下的 world_bank.csv,并把第 0 列(国家名)设为 DataFrame 的索引(index);
  • wb.head():显示前 5 行,快速检查数据长什么样。
ContinentCountryPrimary completion rate: Male: 2015…
AfricaAlgeria106.0…
AfricaAngolaNaN…
AfricaBenin83.0…
AfricaBotswana98.0…
AfricaBurundi58.0…

Continent 是分类变量(qualitative),GDP 增长率、人均国民收入等是定量变量(quantitative)——这决定了下文该画什么图。

1. 定性变量:柱状图(Bar Plot)的三种画法

Cell 4:pandas 自带绘图

wb['Continent'].value_counts().plot(kind='bar');
  • wb['Continent']:取出 Continent 这一列(Series);
  • .value_counts():统计每个大洲出现的次数,返回一个按频数降序排列的 Series;
  • .plot(kind='bar'):调用 pandas 内置绘图,kind='bar' 指定画柱状图;
  • 结尾分号 ;:抑制 Jupyter 输出 <Figure size ...> 这类文本,只显示图。

pandas 柱状图

可以看到非洲的柱最高、南美洲最短。pandas 绘图虽然一行搞定,但功能有限,Data 100 不作为首选。

Cell 6:matplotlib 手写

import matplotlib.pyplot as plt # matplotlib is typically given the alias plt
 
continent = wb['Continent'].value_counts()
plt.bar(continent.index, continent)
plt.xlabel('Continent')
plt.ylabel('Count');
  • import matplotlib.pyplot as plt:导入 matplotlib 的 pyplot 接口,惯例别名 plt;
  • continent = wb['Continent'].value_counts():把频数 Series 存到变量 continent;
  • plt.bar(continent.index, continent):画柱状图——x 轴是类别(continent.index),y 轴是频数(continent);
  • plt.xlabel('Continent') / plt.ylabel('Count'):手动加轴标签。matplotlib 不会自动标注坐标轴,必须自己写;
  • 分号同理,抑制多余文本输出。

matplotlib 柱状图

Cell 8:seaborn 一行搞定

import seaborn as sns # seaborn is typically given the alias sns
sns.countplot(data = wb, x = 'Continent', hue='Continent');
  • import seaborn as sns:导入 seaborn,别名 sns;
  • sns.countplot(data=wb, x='Continent'):直接传入整个 DataFrame 和要统计的列名,seaborn 自动计数并画图;
  • hue='Continent':按大洲给柱子着色(这里颜色只起区分作用,不编码额外信息)。

seaborn 柱状图

seaborn 底层就是 matplotlib,只是接口更简洁,会自动加轴标签和图例。

2. 定量变量不能画柱状图:overplotting

Cell 10:再看一眼数据

wb.head(5)
  • wb.head(5):显示前 5 行。目的是思考:如果要把 Gross national income per capita 当”类别”,每个数值都会是一根柱。

Cell 12:硬用 countplot 的结果

sns.countplot(data = wb, x = 'Gross national income per capita, Atlas method: $: 2016');
  • sns.countplot(data=wb, x='Gross national income per capita, Atlas method: $: 2016'):对定量变量的每一个唯一数值画一根柱,结果就是下面这张几乎无法解读的图。

overplotting

绝大多数柱高为 1(每个值只出现一次),看不出任何信息。结论:定量变量用 histogram / box plot / violin plot。

3. 直方图(Histogram)

Cell 14:matplotlib 版

gni = wb["Gross national income per capita, Atlas method: $: 2016"]
plt.hist(gni, density=True, edgecolor="white")
 
# Add labels
plt.xlabel("Gross national income per capita")
plt.ylabel("Density")
plt.title("Distribution of gross national income per capita");
  • gni = wb["Gross national income per capita, Atlas method: $: 2016"]:把这一长名列取出存为 gni(后面多次复用);
  • plt.hist(gni, density=True, edgecolor="white"):画直方图。density=True 是关键——让纵轴变成 density,每个 bin 的面积与数据比例成正比;edgecolor="white" 给柱子描白边;
  • plt.xlabel(...) / plt.ylabel("Density") / plt.title(...):轴标签和标题,都是手动加的。

matplotlib 直方图

Cell 15:seaborn 版

sns.histplot(data=wb, x="Gross national income per capita, Atlas method: $: 2016", stat="density")
plt.title("Distribution of gross national income per capita");
  • sns.histplot(data=wb, x="...", stat="density"):seaborn 直方图。stat="density" 等价于上面的 density=True,面积与比例成正比;
  • plt.title(...):加标题(seaborn 不负责标题)。

seaborn 直方图

4. 叠加直方图:用 hue 比较类别

Cell 17:先造一个新列 Hemisphere

# Create a new variable to store the hemisphere in which each country is located
north = ["Asia", "Europe", "N. America"]
south = ["Africa", "Oceania", "S. America"]
wb.loc[wb["Continent"].isin(north), "Hemisphere"] = "Northern"
wb.loc[wb["Continent"].isin(south), "Hemisphere"] = "Southern"
  • north = ["Asia", "Europe", "N. America"]:定义北半球大洲列表;
  • south = ["Africa", "Oceania", "S. America"]:定义南半球大洲列表;
  • wb.loc[wb["Continent"].isin(north), "Hemisphere"] = "Northern":
    • wb["Continent"].isin(north) 返回布尔 Series(该行大洲是否在北半球列表);
    • wb.loc[布尔掩码, "Hemisphere"] 选中满足条件的行的 Hemisphere 列;
    • 赋值 "Northern" 即新建/填充该列;
  • 最后一行同理,把南半球国家标为 "Southern"。

这个 cell 没有输出,只是为下一步准备数据。

Cell 18:hue 叠加

sns.histplot(data=wb, x="Gross national income per capita, Atlas method: $: 2016", hue="Hemisphere", stat="density")
plt.title("Distribution of gross national income per capita");
  • hue="Hemisphere":按半球分颜色,两张密度分布叠在一张图里,方便同 bin 比较;
  • seaborn 自动生成图例(legend)——只要颜色编码了信息,就必须有图例;
  • stat="density":仍然是密度,面积=比例。

叠加直方图

5. 验证:“面积 = 比例”

Cell 20

densities, bins, _ = plt.hist(gni, density=True, edgecolor="white", bins=5)
plt.xlabel("Gross national income per capita")
plt.ylabel("Density")
 
print(f"First bin has width {bins[1]-bins[0]} and height {densities[0]}")
print(f"This corresponds to {bins[1]-bins[0]} * {densities[0]} = {(bins[1]-bins[0])*densities[0]*100}% of the data")
  • densities, bins, _ = plt.hist(gni, density=True, edgecolor="white", bins=5):plt.hist 返回三个值——每个 bin 的密度数组 densities、边界数组 bins、画出的 patches(用 _ 丢弃);bins=5 表示分成 5 个箱;
  • plt.xlabel(...) / plt.ylabel("Density"):轴标签;
  • 第一个 print:打印第一箱的宽度()和高度(密度 );
  • 第二个 print:验证 宽 × 高 × 100 = ,即第一箱的面积恰好等于落在其中的数据百分比——这就是”直方图面积=比例,而不是高度”的实证。

运行输出:

First bin has width 16410.0 and height 4.7741589911386953e-05
This corresponds to 16410.0 * 4.7741589911386953e-05 = 78.343949044586% of the data

面积=比例验证

6. 解读直方图:偏度、峰数、离群值

Cell 22:右偏(positive skew)

sns.histplot(data = wb, x = 'Gross national income per capita, Atlas method: $: 2016', stat = 'density');
plt.title('Distribution with a long right tail');
  • sns.histplot(..., stat='density'):画 gni 的密度直方图;
  • plt.title('Distribution with a long right tail'):标题提示这是长右尾分布。

右偏

长尾在右 → 右偏 / 正偏,少数大的离群值把 mean 拉到 median 右边。

Cell 24:左偏(negative skew)

sns.histplot(data = wb, x = 'Access to an improved water source: % of population: 2015', stat = 'density');
plt.title('Distribution with a long left tail');
  • x 换成”改善水源覆盖率”列,其他参数与上面相同;标题提示长左尾。

左偏

Cell 26 / 27 / 29:bins 数量影响峰数判断

# Cell 26:重命名长列名,5 个 bin
wb = wb.rename(columns={'Antiretroviral therapy coverage: % of people living with HIV: 2015':"HIV rate"})
sns.histplot(data=wb, x="HIV rate", stat="density", bins=5)
plt.title("5 histogram bins");
# Cell 27:10 个 bin
sns.histplot(data=wb, x="HIV rate", stat="density", bins=10)
plt.title("10 histogram bins");
# Cell 29:20 个 bin
sns.histplot(data=wb, x ="HIV rate", stat="density", bins=20)
plt.title("20 histogram bins");
  • wb.rename(columns={'...': "HIV rate"}):把超长列名改成 HIV rate(返回新 DataFrame,必须重新赋值给 wb);
  • 三个 cell 只改 bins 参数:5、10、20。注意同一个分布,bin 数不同,看到的峰数不同——这是直方图解读中必须警惕的主观性。

5 bins 10 bins 20 bins

5 bins 看起来单峰(unimodal),10 bins 双峰(bimodal),20 bins 已经很难说哪里算峰。这种模糊性正是后面引入 KDE 的动机之一。

Cell 31:四分位数与中间 50%

gdp = wb['Gross domestic product: % growth : 2016']
gdp = gdp[~gdp.isna()]
 
q1, q2, q3 = np.percentile(gdp, [25, 50, 75])
 
wb_quartiles = wb.copy()
wb_quartiles['category'] = None
wb_quartiles.loc[(wb_quartiles['Gross domestic product: % growth : 2016'] < q1) | (wb_quartiles['Gross domestic product: % growth : 2016'] > q3), 'category'] = 'Outside of the middle 50%'
wb_quartiles.loc[(wb_quartiles['Gross domestic product: % growth : 2016'] > q1) & (wb_quartiles['Gross domestic product: % growth : 2016'] < q3), 'category'] = 'In the middle 50%'
 
sns.histplot(wb_quartiles, x="Gross domestic product: % growth : 2016", hue="category")
sns.rugplot([q1, q2, q3], c="firebrick", lw=6, height=0.1);
  • gdp = wb['Gross domestic product: % growth : 2016']:取 GDP 增长率列;
  • gdp = gdp[~gdp.isna()]:gdp.isna() 得到缺失值掩码,~ 取反后筛掉 NaN;
  • q1, q2, q3 = np.percentile(gdp, [25, 50, 75]):一次算出 25th / 50th(median)/ 75th 百分位数;
  • wb_quartiles = wb.copy():先复制 DataFrame,避免修改原数据;
  • wb_quartiles['category'] = None:新建 category 列并初始化为 None;
  • 第一行 loc[...] = 'Outside of the middle 50%':条件 (gdp < q1) | (gdp > q3)(注意 | 是 or,必须加括号)——落在中间 50% 之外的行标上 Outside...;
  • 第二行 loc[...] = 'In the middle 50%':条件 (gdp > q1) & (gdp < q3) 标中间 50%;
  • sns.histplot(wb_quartiles, x=..., hue="category"):按”中间/外部”着色画直方图,直观看到中间 50% 的范围;
  • sns.rugplot([q1, q2, q3], c="firebrick", lw=6, height=0.1):在 Q1/Q2/Q3 位置画红色粗短线(rug plot),标出四分位数位置。

四分位数

离群值定义: 或 (IQR = Q3 − Q1)。


后半部分(箱线图、小提琴图、KDE)见 2026-08-06-Note。