實作流程
Step by step:五個階段
bash
# 確認 ollama 正在執行
$ ollama --version
ollama version is 0.5.x
# 下載嵌入模型(274 MB,比 LLM 小很多)
$ ollama pull nomic-embed-text
pulling manifest ...
pulling 970aa74c0a90... ████████████████████ 100% 274 MB
success
# 確認已下載
$ ollama list
NAME SIZE MODIFIED
nomic-embed-text 274 MB just now
bash
$ pip install requests scikit-learn matplotlib --break-system-packages
Successfully installed ...
Python — get_embeddings.py
import requests, json
# ── 你的詞彙表(自由修改!)──────────────────────
TW_WORDS = ["軟體", "硬體", "滑鼠", "影片", "計程車", "公車",
"馬鈴薯", "泡麵", "網紅", "直播主", "打卡", "本土"]
CN_WORDS = ["軟件", "硬件", "鼠標", "視頻", "出租車", "公交車",
"土豆", "方便麵", "博主", "主播", "打卡", "本土"]
# 注意:「打卡」和「本土」是同形詞,各出現一次
# ───────────────────────────────────────────────
def get_embedding(word):
resp = requests.post("http://localhost:11434/api/embed",
json={"model": "nomic-embed-text", "input": word})
return resp.json()["embeddings"][0]
print("取得詞向量中...")
tw_vecs = [get_embedding(w) for w in TW_WORDS]
cn_vecs = [get_embedding(w) for w in CN_WORDS]
print(f"完成!每個向量維度:{len(tw_vecs[0])}")
# 存起來,下一步用
data = {"tw": {"words": TW_WORDS, "vecs": tw_vecs},
"cn": {"words": CN_WORDS, "vecs": cn_vecs}}
json.dump(data, open("embeddings.json","w"), ensure_ascii=False)
print("已儲存至 embeddings.json")
Python — visualize.py
import json, numpy as np
import matplotlib.pyplot as plt
import matplotlib.font_manager as fm
from sklearn.manifold import TSNE
# ── 讀取向量 ────────────────────────────────────
data = json.load(open("embeddings.json"))
tw_words = data["tw"]["words"]
cn_words = data["cn"]["words"]
all_vecs = np.array(data["tw"]["vecs"] + data["cn"]["vecs"])
labels = tw_words + cn_words
colors = ["#1a4a9e"] * len(tw_words) + ["#8b1a1a"] * len(cn_words)
# ── t-SNE 降維到 2D ──────────────────────────────
tsne = TSNE(n_components=2, random_state=42, perplexity=5)
coords = tsne.fit_transform(all_vecs)
# ── 畫圖 ────────────────────────────────────────
fig, ax = plt.subplots(figsize=(12, 9))
ax.set_facecolor("#faf9f7")
fig.patch.set_facecolor("#faf9f7")
for i, (x, y) in enumerate(coords):
word = labels[i]
c = colors[i]
is_shared = (word in tw_words and word in cn_words)
marker = "D" if is_shared else "o" # 菱形 = 同形詞
ax.scatter(x, y, c=c, s=120 if is_shared else 80,
marker=marker, zorder=3, alpha=0.85)
ax.annotate(word, (x, y),
textcoords="offset points", xytext=(6, 4),
fontsize=11, color=c, fontweight="bold")
# ── 畫連線:台灣詞 ↔ 對應中國詞 ────────────────
for i, tw in enumerate(tw_words):
if tw in cn_words:
continue # 同形詞跳過
if i < len(cn_words):
j = len(tw_words) + i # 對應中國詞在 all_vecs 的索引
ax.plot([coords[i,0], coords[j,0]],
[coords[i,1], coords[j,1]],
color="#aaa", linewidth=0.6, linestyle="--", zorder=1)
# ── 圖例與標題 ──────────────────────────────────
from matplotlib.lines import Line2D
legend = [
Line2D([0],[0],marker='o',color='w',markerfacecolor='#1a4a9e',ms=9,label='台灣漢語'),
Line2D([0],[0],marker='o',color='w',markerfacecolor='#8b1a1a',ms=9,label='中國漢語'),
Line2D([0],[0],marker='D',color='w',markerfacecolor='#555',ms=9,label='同形詞'),
]
ax.legend(handles=legend, loc='upper left', framealpha=.8, fontsize=11)
ax.set_title("兩岸漢語詞彙的向量空間分布\n(nomic-embed-text · t-SNE 降維)",
fontsize=14, pad=14)
ax.axis("off")
plt.tight_layout()
plt.savefig("lexical_map.png", dpi=150, bbox_inches="tight")
print("圖片已儲存為 lexical_map.png")
plt.show()
中文字型問題:matplotlib 預設不支援中文。macOS 可在程式頂端加 plt.rcParams['font.family'] = 'Heiti TC';Windows 加 'Microsoft JhengHei';Linux 需先安裝 Noto Sans CJK。
Python — similarity.py
import json, numpy as np
def cosine(a, b):
a, b = np.array(a), np.array(b)
return float(np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b)))
data = json.load(open("embeddings.json"))
tw_w, tw_v = data["tw"]["words"], data["tw"]["vecs"]
cn_w, cn_v = data["cn"]["words"], data["cn"]["vecs"]
print(f"{'台灣詞':<8} {'中國詞':<8} {'餘弦相似度':>8} 解讀")
print("─" * 52)
for i in range(min(len(tw_w), len(cn_w))):
sim = cosine(tw_v[i], cn_v[i])
note = "高度相近" if sim > 0.9 else ("相近" if sim > 0.75 else "有語義落差")
print(f"{tw_w[i]:<8} {cn_w[i]:<8} {sim:>8.4f} {note}")
預期輸出(示意)
台灣詞 中國詞 餘弦相似度 解讀
────────────────────────────────────────────────────
軟體 軟件 0.9312 高度相近
硬體 硬件 0.9287 高度相近
滑鼠 鼠標 0.8841 相近
影片 視頻 0.8654 相近
計程車 出租車 0.9156 高度相近
馬鈴薯 土豆 0.8723 相近
打卡 打卡 0.9998 高度相近 ← 同形詞與自身比較
本土 本土 0.9998 高度相近 ← 嵌入模型無法區分兩岸語境
有趣的觀察:同形詞(如「打卡」「本土」)與自身比較的相似度接近 1.0,代表嵌入模型把它們當成同一個詞——這本身就是一個值得討論的發現:模型學到了「字形」,但未必學到「語境差異」。這是你反思作業的好素材。