Word2Vec词向量:让机器理解词语的语义
2026年
先验直觉:在自然语言处理中,最基础的问题莫过于:如何让计算机理解词语的含义? 传统方法使用 One-Hot 编码将每个词表示为一个高维稀疏向量(向量的维度等于词典大小,该词对应位置为 1,其余位置为 0)。这种方法有两个致命缺陷:
关键词:Python,scikit-learn,matplotlib,BERT,PCA,Embedding,过拟合,聚类
- 引言
"You shall know a word by the company it keeps." — J.R. Firth (1957)
在自然语言处理中,最基础的问题莫过于:如何让计算机理解词语的含义? 传统方法使用 One-Hot 编码将每个词表示为一个高维稀疏向量(向量的维度等于词典大小,该词对应位置为 1,其余位置为 0)。这种方法有两个致命缺陷:
- 维度灾难:词典大小动辄数十万,每个词向量的维度也高达数十万,存储和计算成本极高。
- 语义鸿沟:任意两个不同的词的 One-Hot 向量相互正交(内积为 0),
,完全无法体现语义上的相似性——"猫"和"狗"的距离与"猫"和"冰箱"的距离完全相同。
分布式表示(Distributed Representation) 的出现彻底改变了这一局面。其核心思想是:用稠密的低维向量来表示每个词,向量的每个维度不再对应于词典中的一个具体词,而是编码了该词的某种潜在语义特征。语义相近的词,其向量在欧几里得空间中也彼此靠近。
Word2Vec 由 Google 的 Mikolov 等人在 2013 年提出,是分布式表示方法中最具影响力的模型之一。它将每个词映射到一个 50~300 维的稠密向量中,使得语义相近的词在向量空间中彼此靠近,甚至支持令人惊叹的向量代数运算:
本文将系统地从理论基础到代码实践,完整拆解 Word2Vec 的 CBOW 和 Skip-gram 两种架构,并动手训练词向量、计算词语相似度、进行向量代数运算和可视化分析。通过本文的学习,你将彻底理解词向量的工作原理,并掌握用 gensim 工具完成词嵌入任务的完整流程。
- 理论基础
2.1 分布式假设:一词之友可见其意
分布式假设(Distributional Hypothesis) 是 Word2Vec 的哲学基础:上下文相似的词,其语义也相似。
例如,观察以下句子:
- "我爱吃苹果"
- "我爱吃香蕉"
- "我爱吃米饭"
"苹果""香蕉""米饭"都出现在"我爱吃____"的语境中,因此它们的语义在某种程度上是相近的(都是食物)。Word2Vec 正是通过预测上下文或被上下文预测来学习词向量。
2.2 CBOW(Continuous Bag of Words)
CBOW 的目标:给定上下文词(窗口中除中心词外的邻居词),预测中心词。形象地说,这就像"完形填空"——给你前后文,猜中间被遮住的那个词。
模型结构
输入层(上下文词) 投影层(求和平均) 输出层(Softmax)
w_{t-2} → v_{t-2} ─┐
w_{t-1} → v_{t-1} ─┼→ h = avg(v) ──→ softmax(u^T h) ──→ p(w_t | context)
w_{t+1} → v_{t+1} ─┤
w_{t+2} → v_{t+2} ─┘设窗口大小为
数学推导
第一步:每个词
- 输入向量
:当该词作为上下文词时使用的向量,即输入层到投影层的权重矩阵 的第 行。 - 输出向量
:当该词作为中心词时使用的向量,即投影层到输出层的权重矩阵 的第 列。
训练完成后,通常取输入向量
第二步:将上下文词(共
这里取平均而非拼接,保证了上下文顺序的不变性(Bag of Words 的含义),也使得输入维度固定。
第三步:计算每个词
其中
第四步:优化目标是最大化整个语料的对数似然(每个窗口生成一个训练样本):
使用随机梯度下降(SGD)更新参数。对每个训练样本
直观理解:输出向量
直观理解
CBOW 像"完形填空":给你"我爱____苹果",你猜空白处最可能是"吃"。
2.3 Skip-gram
Skip-gram 的目标:给定中心词,预测上下文词。与 CBOW 相反。
模型结构
输入层(中心词) → 投影层 → 输出层(Softmax预测多个上下文词)数学推导
第一步:给定中心词
第二步:对每个上下文位置
这意味着对于每个中心词
第三步:整个窗口的对数似然为所有位置预测概率的对数之和:
与 CBOW 相比,Skip-gram 对每个中心词生成
CBOW vs Skip-gram 对比
| 特性 | CBOW | Skip-gram |
|---|---|---|
| 预测方向 | 上下文 → 中心词 | 中心词 → 上下文 |
| 训练速度 | 更快(一次更新所有上下文) | 较慢(每个中心词更新多个样本) |
| 低频词表现 | 一般 | 更好(低频词被多次采样) |
| 适用场景 | 大规模语料、高频词 | 小规模语料、低频词重要 |
2.4 优化技巧
直接对全词表做 Softmax 的计算复杂度为
负采样(Negative Sampling, NEG)
核心思想:不再计算昂贵的全词表 Softmax,而是将多分类问题巧妙地转化为二分类问题——判别一个词对
具体做法:
- 正样本:从语料中取真实存在的中心词-上下文词对
,标签为 1。 - 负样本:随机从词表中采样
个词作为噪声词,与中心词组成"假"词对 ,标签为 0。
Sigmoid 二分类的目标函数如下:
其中:
是 Sigmoid 函数- 第一项最大化正样本的匹配得分
- 第二项最小化负样本的匹配得分(即最大化
的 Sigmoid)
噪声分布
这个 0.75 指数的作用是:降低高频词的采样概率,提高低频词被选为负样本的机会,从而让模型更均匀地学习所有词的表示。
负采样的优势:每次参数更新只涉及
层次 Softmax(Hierarchical Softmax)
层次 Softmax 是另一种避开全词表 Softmax 的方法,它利用霍夫曼树(Huffman Tree) 对词表进行编码。具体来说:
- 按照词频构建霍夫曼树:高频词路径短,低频词路径长。
- 每个叶子节点对应一个词,内部节点对应一个二分类器(一个可学习的向量
)。 - 从根节点走到目标词
的叶子节点,路径上经过 个内部节点,每个节点做一次二分类决策(向左走 vs 向右走)。
数学表达如下:
其中
层次 Softmax vs 负采样的选择建议:
- 层次 Softmax:适合词典极大(
)且低频词多的场景,因为树结构天然利用了词频信息。 - 负采样:适合中等规模语料,训练速度更快,词向量质量通常更好(尤其在语义相似度任务上)。
- 环境准备
python
# 安装所需库
# pip install gensim numpy matplotlib scikit-learn
import numpy as np
import matplotlib.pyplot as plt
from sklearn.decomposition import PCA
from sklearn.manifold import TSNE
import gensim
from gensim.models import Word2Vec
import warnings
warnings.filterwarnings('ignore')
print("所有库导入成功")
# 预期输出: 所有库导入成功- 用 gensim 训练 Word2Vec
4.1 准备小语料
我们使用 text8 语料(Wikipedia 前 10^8 字符的预处理版本),体积小、质量高。
python
import urllib.request
import os
# 下载 text8 语料(约 31MB)
url = "https://raw.githubusercontent.com/dwyl/english-words/master/words_alpha.txt"
# 改用内置的简单语料:直接用几段英文文本构建
corpus_text = [
"the king is the ruler of the kingdom",
"the queen is the ruler of the kingdom",
"the man is strong and powerful",
"the woman is beautiful and kind",
"the king and queen rule the kingdom together",
"the boy plays in the park",
"the girl plays in the park",
"the cat chases the mouse",
"the dog chases the cat",
"the mouse hides in the hole",
"the king wears a crown",
"the queen wears a crown",
"the man works in the office",
"the woman works in the office",
"the boy studies at school",
"the girl studies at school",
"the cat sleeps on the sofa",
"the dog sleeps on the bed",
"paris is the capital of france",
"london is the capital of england",
"berlin is the capital of germany",
"china is a large country in asia",
"japan is a country in asia",
"france is a country in europe",
"germany is a country in europe",
"the river flows through the valley",
"the mountain is covered with snow",
"the sun rises in the east",
"the moon shines at night",
"the star twinkles in the sky",
]
# 分词(按空格拆分)
sentences = [sentence.lower().split() for sentence in corpus_text]
print(f"语料句子数: {len(sentences)}")
print(f"前两个句子: {sentences[:2]}")
print(f"总词数: {sum(len(s) for s in sentences)}")
# 预期输出:
# 语料句子数: 30
# 前两个句子: [['the', 'king', 'is', 'the', 'ruler', 'of', 'the', 'kingdom'], ...]
# 总词数: 2024.2 训练 CBOW 模型
python
# 训练 CBOW 模型
# sg=0 表示使用 CBOW 架构
# vector_size=100:词向量维度
# window=5:上下文窗口大小
# min_count=1:忽略词频小于1的词(保留全部)
# workers=4:并行线程数
model_cbow = Word2Vec(
sentences=sentences,
vector_size=100,
window=5,
min_count=1,
sg=0, # 0=CBOW, 1=Skip-gram
negative=5, # 负采样个数
epochs=100, # 迭代次数
workers=4
)
print(f"词典大小: {len(model_cbow.wv)}")
print(f"词向量维度: {model_cbow.wv.vector_size}")
print(f"'king' 的词向量 (前10维): {model_cbow.wv['king'][:10]}")
# 预期输出:
# 词典大小: 68
# 词向量维度: 100
# 'king' 的词向量 (前10维): [-0.123 0.456 -0.789 ...] (随机初始化的具体值)4.3 训练 Skip-gram 模型
python
# 训练 Skip-gram 模型
model_sg = Word2Vec(
sentences=sentences,
vector_size=100,
window=5,
min_count=1,
sg=1, # 1=Skip-gram
negative=5,
epochs=100,
workers=4
)
print(f"Skip-gram 词典大小: {len(model_sg.wv)}")
print(f"Skip-gram 'king' 向量 (前10维): {model_sg.wv['king'][:10]}")
# 预期输出:
# Skip-gram 词典大小: 68
# Skip-gram 'king' 向量 (前10维): [ ... ]- 词向量相似度分析
5.1 计算最相似词(most_similar)
python
# 找出与 'king' 最相似的词
similar_to_king = model_cbow.wv.most_similar('king', topn=5)
print("与 'king' 最相似的词 (CBOW):")
for word, score in similar_to_king:
print(f" {word:12s} 相似度: {score:.4f}")
# 预期输出 (示例):
# 与 'king' 最相似的词 (CBOW):
# queen 相似度: 0.9234
# ruler 相似度: 0.8567
# kingdom 相似度: 0.8123
# crown 相似度: 0.7845
# man 相似度: 0.6543python
# 对比 Skip-gram 的结果
similar_to_king_sg = model_sg.wv.most_similar('king', topn=5)
print("与 'king' 最相似的词 (Skip-gram):")
for word, score in similar_to_king_sg:
print(f" {word:12s} 相似度: {score:.4f}")
# 预期输出: 结果与 CBOW 类似但排序和得分略有差异python
# 查找不相似的词(doesnt_match)
odd_word = model_cbow.wv.doesnt_match(["king", "queen", "crown", "dog"])
print(f"不匹配的词: {odd_word}")
# 预期输出: 不匹配的词: dog5.2 最相似词条形图可视化
python
def plot_most_similar(model, word, topn=8):
"""绘制与目标词最相似的 topn 个词及相似度"""
similar = model.wv.most_similar(word, topn=topn)
words, scores = zip(*similar)
fig, ax = plt.subplots(figsize=(10, 6))
colors = plt.cm.Blues(np.linspace(0.4, 0.9, topn))
bars = ax.barh(range(topn), scores, color=colors[::-1])
ax.set_yticks(range(topn))
ax.set_yticklabels(words)
ax.invert_yaxis()
ax.set_xlabel('Cosine Similarity', fontsize=12)
ax.set_title(f"Top {topn} Words Most Similar to '{word}'", fontsize=14, fontweight='bold')
# 在条形上添加数值标签
for i, (bar, score) in enumerate(zip(bars, scores)):
ax.text(score + 0.01, bar.get_y() + bar.get_height()/2,
f'{score:.3f}', va='center', fontsize=10)
plt.tight_layout()
plt.savefig('/home/wq/projects/qian-style-articles/articles/images/word2vec_similarity_bar.png', dpi=150)
plt.show()
print(f"可视化已保存")
plot_most_similar(model_cbow, 'king')(图:Word2Vec 相似度条形图 - 需运行 gen_figures.py 生成)
- 词向量的代数运算
Word2Vec 最惊艳的特性之一:向量空间中存在语义方向的线性平移。
6.1 经典例子:King - Man + Woman = ?
python
# 词向量代数运算
def word_algebra(model, positive, negative, topn=5):
"""
向量代数运算: positive中的词相加,negative中的词相减
返回最接近结果向量的词
"""
result = model.wv.most_similar(
positive=positive,
negative=negative,
topn=topn
)
return result
# 经典运算: king - man + woman ≈ queen
result_1 = word_algebra(model_cbow,
positive=['king', 'woman'],
negative=['man'])
print("king - man + woman = ?")
for word, score in result_1:
print(f" {word:12s} 得分: {score:.4f}")
# 预期输出 (示例):
# king - man + woman = ?
# queen 得分: 0.8145
# ruler 得分: 0.7234
# kingdom 得分: 0.6842
# crown 得分: 0.6541
# woman 得分: 0.6123python
# 国家-首都关系: Paris - France + Germany ≈ Berlin
result_2 = word_algebra(model_cbow,
positive=['paris', 'germany'],
negative=['france'])
print("paris - france + germany = ?")
for word, score in result_2:
print(f" {word:12s} 得分: {score:.4f}")
# 预期输出:
# paris - france + germany = ?
# berlin 得分: 0.7243
# germany 得分: 0.6987
# ...python
# 更多代数运算
# 男性-女性关系对
result_3 = word_algebra(model_cbow,
positive=['king', 'girl'],
negative=['queen'])
print("king + girl - queen = ?")
for word, score in result_3:
print(f" {word:12s} 得分: {score:.4f}")
# 预期输出: boy 应出现在前列6.2 代数运算结果展示表
python
# 批量测试多个代数关系并展示为表格
relations = [
(['king', 'woman'], ['man'], '国王-男人+女人'),
(['paris', 'germany'], ['france'], '巴黎-法国+德国'),
(['london', 'china'], ['england'], '伦敦-英国+中国'),
(['man', 'girl'], ['woman'], '男人+女孩-女人'),
(['france', 'berlin'], ['germany'], '法国+柏林-德国'),
]
print("=" * 70)
print(f"{'代数运算':^30s} | {'预期结果':^12s} | {'Top1':^12s} | {'相似度':^8s}")
print("=" * 70)
for pos, neg, desc in relations:
result = word_algebra(model_cbow, positive=pos, negative=neg, topn=3)
top_word = result[0][0] if result else "N/A"
top_score = result[0][1] if result else 0
# 根据运算类型推断预期结果
expected = {
'国王-男人+女人': 'queen',
'巴黎-法国+德国': 'berlin',
'伦敦-英国+中国': 'china',
'男人+女孩-女人': 'boy',
'法国+柏林-德国': 'paris',
}.get(desc, '')
print(f"{desc:^28s} | {expected:^10s} | {top_word:^10s} | {top_score:.4f}")
print("=" * 70)
# 预期输出: 表格形式的代数运算结果- 词向量空间可视化:PCA 与 t-SNE
200 维以上的向量无法直接观察,需要降维到 2D 进行可视化。
7.1 PCA 降维 + 语义染色
python
# 选择一组有语义类别的词
word_categories = {
'ROYALTY': ['king', 'queen', 'prince', 'crown', 'ruler'],
'GENDER': ['man', 'woman', 'boy', 'girl'],
'ANIMALS': ['cat', 'dog', 'mouse'],
'COUNTRIES': ['france', 'germany', 'china', 'japan', 'england'],
'GEOGRAPHY': ['river', 'mountain', 'valley', 'sea', 'snow'],
'ASTRONOMY': ['sun', 'moon', 'star', 'sky'],
}
# 收集所有词及其类别标签
all_words = []
word_labels = []
colors = []
color_map = {
'ROYALTY': 'red', 'GENDER': 'blue', 'ANIMALS': 'green',
'COUNTRIES': 'orange', 'GEOGRAPHY': 'purple', 'ASTRONOMY': 'cyan',
}
for category, words in word_categories.items():
for w in words:
if w in model_cbow.wv:
all_words.append(w)
word_labels.append(category)
colors.append(color_map[category])
print(f"待可视化词数: {len(all_words)}")
print(f"词及其类别: {list(zip(all_words, word_labels))}")
# 预期输出: 待可视化词数: ~20个词及其类别python
# 提取词向量
word_vectors = np.array([model_cbow.wv[w] for w in all_words])
print(f"词向量矩阵形状: {word_vectors.shape}")
# 预期输出: 词向量矩阵形状: (20, 100)
# PCA 降维到 2D
pca = PCA(n_components=2)
vectors_pca = pca.fit_transform(word_vectors)
print(f"PCA 降维后形状: {vectors_pca.shape}")
print(f"前两个主成分解释方差比: {pca.explained_variance_ratio_}")
# 预期输出: PCA 降维后形状: (20, 2)
# 前两个主成分解释方差比: [0.312 0.184]python
# 绘制 PCA 散点图
fig, ax = plt.subplots(figsize=(12, 9))
# 按类别绘制不同颜色的点
unique_categories = list(color_map.keys())
for cat in unique_categories:
mask = [label == cat for label in word_labels]
if sum(mask) > 0:
ax.scatter(vectors_pca[mask, 0], vectors_pca[mask, 1],
c=color_map[cat], label=cat, s=120, alpha=0.8, edgecolors='black', linewidth=0.5)
# 添加词标签
for i, word in enumerate(all_words):
ax.annotate(word, (vectors_pca[i, 0], vectors_pca[i, 1]),
fontsize=11, fontweight='bold',
xytext=(5, 5), textcoords='offset points')
ax.set_title('Word2Vec 词向量 PCA 降维可视化 (按语义染色)', fontsize=16, fontweight='bold')
ax.set_xlabel('PC1', fontsize=12)
ax.set_ylabel('PC2', fontsize=12)
ax.legend(fontsize=11, loc='best')
ax.grid(True, alpha=0.3)
plt.tight_layout()
plt.savefig('/home/wq/projects/qian-style-articles/articles/images/word2vec_pca.png', dpi=150)
plt.show()
print("PCA 可视化已保存")(图:Word2Vec PCA 可视化 - 需运行 gen_figures.py 生成)
7.2 t-SNE 降维可视化
python
# t-SNE 降维(perplexity 控制局部 vs 全局结构)
tsne = TSNE(n_components=2, perplexity=5, random_state=42, n_iter=1000)
vectors_tsne = tsne.fit_transform(word_vectors)
print(f"t-SNE 降维后形状: {vectors_tsne.shape}")
# 预期输出: t-SNE 降维后形状: (20, 2)python
# 绘制 t-SNE 散点图
fig, ax = plt.subplots(figsize=(12, 9))
for cat in unique_categories:
mask = [label == cat for label in word_labels]
if sum(mask) > 0:
ax.scatter(vectors_tsne[mask, 0], vectors_tsne[mask, 1],
c=color_map[cat], label=cat, s=120, alpha=0.8, edgecolors='black', linewidth=0.5)
for i, word in enumerate(all_words):
ax.annotate(word, (vectors_tsne[i, 0], vectors_tsne[i, 1]),
fontsize=11, fontweight='bold',
xytext=(5, 5), textcoords='offset points')
ax.set_title('Word2Vec 词向量 t-SNE 降维可视化 (按语义染色)', fontsize=16, fontweight='bold')
ax.set_xlabel('t-SNE Dimension 1', fontsize=12)
ax.set_ylabel('t-SNE Dimension 2', fontsize=12)
ax.legend(fontsize=11, loc='best')
ax.grid(True, alpha=0.3)
plt.tight_layout()
plt.savefig('/home/wq/projects/qian-style-articles/articles/images/word2vec_tsne.png', dpi=150)
plt.show()
print("t-SNE 可视化已保存")(图:Word2Vec t-SNE 可视化 - 需运行 gen_figures.py 生成)
7.3 PCA vs t-SNE 对比解读
| 特点 | PCA | t-SNE |
|---|---|---|
| 本质 | 线性变换(方差最大化) | 非线性流形学习 |
| 全局结构 | 保留(全局距离关系好) | 可能扭曲(更关注局部邻域) |
| 运行速度 | 快( | 慢( |
| 词向量可视化 | 适合看大类簇的分布 | 适合看局部相似词的聚集 |
| 重复运行 | 结果确定 | 每次结果不同(需固定 seed) |
- CBOW vs Skip-gram 训练速度对比
python
import time
def train_and_time(sentences, sg, epochs=200, label=''):
"""训练并计时"""
start = time.time()
model = Word2Vec(
sentences=sentences,
vector_size=100,
window=5,
min_count=1,
sg=sg,
negative=5,
epochs=epochs,
workers=4
)
elapsed = time.time() - start
return model, elapsed
# 分别训练并计时
print("正在训练 CBOW...")
model_cbow_t, time_cbow = train_and_time(sentences, sg=0, epochs=200, label='CBOW')
print("正在训练 Skip-gram...")
model_sg_t, time_sg = train_and_time(sentences, sg=1, epochs=200, label='Skip-gram')
# 比较
print(f"\n{'='*45}")
print(f"{'模型':^20s} | {'训练时间':^20s}")
print(f"{'='*45}")
print(f"{'CBOW':^20s} | {time_cbow:^20.4f} 秒")
print(f"{'Skip-gram':^20s} | {time_sg:^20.4f} 秒")
print(f"{'='*45}")
print(f"加速比 (Skip-gram / CBOW): {time_sg / time_cbow:.2f}x")
# 预期输出:
# CBOW 通常比 Skip-gram 快 1.5~3 倍python
# 绘制速度对比条形图
fig, ax = plt.subplots(figsize=(8, 5))
models = ['CBOW', 'Skip-gram']
times = [time_cbow, time_sg]
colors = ['#3498db', '#e74c3c']
bars = ax.bar(models, times, color=colors, width=0.5, edgecolor='black', linewidth=1.2)
# 添加数值标签
for bar, t in zip(bars, times):
ax.text(bar.get_x() + bar.get_width()/2, bar.get_height() + 0.05,
f'{t:.2f}s', ha='center', va='bottom', fontsize=14, fontweight='bold')
ax.set_ylabel('Training Time (seconds)', fontsize=13)
ax.set_title('CBOW vs Skip-gram 训练速度对比\n(epochs=200, vector_size=100)',
fontsize=14, fontweight='bold')
ax.set_ylim(0, max(times) * 1.3)
ax.grid(axis='y', alpha=0.3)
plt.tight_layout()
plt.savefig('/home/wq/projects/qian-style-articles/articles/images/word2vec_speed_compare.png', dpi=150)
plt.show()
print("速度对比图已保存")(图:CBOW vs Skip-gram 速度对比 - 需运行 gen_figures.py 生成)
- 完整实验:从训练到可视化的完整 Pipeline
python
def word2vec_pipeline(sentences, vector_size=100, window=5, sg=0, epochs=100):
"""
完整的 Word2Vec 训练 Pipeline
参数:
sentences: 分词后的句子列表
vector_size: 词向量维度
window: 上下文窗口大小
sg: 0=CBOW, 1=Skip-gram
epochs: 迭代次数
返回:
model: 训练好的 Word2Vec 模型
metrics: 训练指标字典
"""
start = time.time()
# Step 1: 训练
model = Word2Vec(
sentences=sentences,
vector_size=vector_size,
window=window,
min_count=1,
sg=sg,
negative=5,
epochs=epochs,
workers=4
)
elapsed = time.time() - start
# Step 2: 收集指标
metrics = {
'vocab_size': len(model.wv),
'vector_size': vector_size,
'window': window,
'model_type': 'CBOW' if sg == 0 else 'Skip-gram',
'train_time': elapsed,
'epochs': epochs,
'total_words': sum(len(s) for s in sentences),
}
# Step 3: 验证语义关系
test_cases = [
(['king', 'woman'], ['man'], 'king - man + woman'),
(['paris', 'germany'], ['france'], 'paris - france + germany'),
]
print(f"\n{'='*50}")
print(f"模型: {metrics['model_type']} | 维度: {vector_size} | 耗时: {elapsed:.2f}s")
print(f"{'='*50}")
for pos, neg, desc in test_cases:
result = model.wv.most_similar(positive=pos, negative=neg, topn=1)
print(f" {desc} ≈ {result[0][0]} ({result[0][1]:.3f})")
return model, metrics
# 运行 Pipeline
model_pipe, metrics = word2vec_pipeline(sentences, vector_size=100, sg=0, epochs=200)
# 输出汇总
print(f"\n{'='*50}")
print("Pipeline 运行汇总:")
for k, v in metrics.items():
print(f" {k}: {v}")
print(f"{'='*50}")
# 预期输出: 完整的训练指标和语义验证结果- 常见问题与调参建议
10.1 参数调优指南
| 参数 | 推荐范围 | 说明 |
|---|---|---|
vector_size | 100~300 | 太小欠拟合,太大过拟合且慢 |
window | 5~10 | 太小缺乏上下文,太大引入噪声 |
min_count | 5~10 | 过滤低频词,节省内存提升质量 |
negative | 5~20 | 越大越慢但可能更准 |
epochs | 5~50 | 语料越大所需 epoch 越少 |
sg | 0 或 1 | 大语料选 CBOW(快),小语料选 Skip-gram(准) |
10.2 常见陷阱
- 未去除停用词:'the', 'a', 'is' 等高频词会淹没语义信号
- 窗口过大:引入无关词作为上下文,降低向量的语义纯度
- epoch 过多:在小语料上反复学习会导致过拟合(向量空间扭曲)
- min_count 过小:罕见词向量质量差,且占用大量内存
python
# 更好的训练配置(实际应用)
model_best = Word2Vec(
sentences=sentences,
vector_size=150,
window=5,
min_count=2, # 过滤低频词
sg=1, # Skip-gram
negative=10, # 更多负样本
ns_exponent=0.75,# 负采样指数
sample=1e-5, # 高频词下采样
epochs=50,
workers=4
)
print(f"优化配置下词典大小: {len(model_best.wv)}")
print(f"相似度检验 - king ≈ queen: {model_best.wv.similarity('king', 'queen'):.4f}")
# 预期输出:
# 优化配置下词典大小: 约 45 (min_count=2 过滤了部分词)
# 相似度检验 - king ≈ queen: 0.89xx- 总结
Word2Vec 通过分布式表示将词语映射到稠密向量空间,是 NLP 领域最具影响力的词嵌入方法之一。本文完整覆盖了:
- 理论基础:分布式假设 → CBOW/Skip-gram 结构 → 数学推导 → 优化技巧
- 代码实现:gensim 训练 → 相似度计算 → 向量代数运算
- 可视化分析:PCA/t-SNE 降维 → 语义聚类散点图 → 相似度条形图 → 速度对比
- 实践建议:参数调优指南与常见陷阱
扩展阅读:
- GloVe(Global Vectors):利用全局共现矩阵的词嵌入
- FastText:引入子词(subword)信息,解决 OOV 问题
- BERT Embeddings:上下文相关的动态词表示
一、数学文化:从分布语义到词向量革命
1.1 约翰·鲁珀特·弗斯(John Rupert Firth, 1890-1960)
英国语言学家,提出了分布语义学(Distributional Semantics)的核心原则:"一个词的意思由它周围的词决定"(You shall know a word by the company it keeps)。这个在1957年提出的语言学洞见,到了深度学习时代被验证为——词向量正是通过上下文预测来学习的。
1.2 托马斯·米科洛夫(Tomas Mikolov, 1982-)
捷克计算机科学家,Google研究员。2013年提出了Word2Vec,包括CBOW(Continuous Bag of Words)和Skip-gram两种模型。米科洛夫的关键创新是用极其简单的单层神经网络(无隐藏层)捕捉词之间的语义关系,训练速度快到可以在单机上处理数十亿词。
1.3 约书亚·本吉奥(Yoshua Bengio, 1964-)
加拿大计算机科学家,图灵奖得主(2018)。他2003年发表的《神经概率语言模型》(Neural Probabilistic Language Model)首次提出用神经网络学习词的分布式表示——即词向量。这篇论文是Word2Vec和所有后续词嵌入工作的理论基础。