From embeddings to labelled maps¶

This notebook turns a corpus of newspaper articles into an interactive map you can explore in the browser. Every dot is an article: articles that are semantically close sit next to each other, and the groups that emerge are named by a language model.

It is the notebook behind the visualisation UMAP of SBA coverage, 1940–1999, built from articles that mention the Swiss Bankers Association in four Swiss newspapers. It works on any corpus with the columns listed below.

The pipeline

# Stage What happens
1 Setup imports, folders, settings, helper functions
2 Embed each article becomes a vector of numbers that encodes its meaning
3 Reduce UMAP squeezes the vectors down to two dimensions, giving every article an (x, y) position
4 First look the map, coloured by year
5 Cluster HDBSCAN finds dense groups of articles on the map
6 Select the most representative articles of each cluster are picked
7 Name a language model proposes a short label for each cluster
8 Label the clusters' names are printed on the map
9 Check tests of whether the labels can be trusted
10 Export a standalone HTML page with a friendly hover box, opacity and size controls, and word search

Requirements. Python 3.10 or later and

pip install impresso umap-learn scikit-learn plotly pandas numpy tqdm ollama

Input. A pickled pandas DataFrame at data/articles.pkl, one row per article, with these columns:

column content
id unique identifier of the article
text.content full text of the article
text.title title of the article (only used when asking the language model to name a cluster)
meta.date publication date, e.g. 1957-03-12
meta.mediaId newspaper code (JDG, GDL, DBB, NZZ)
text.langCode language code (fr, de), only used in the checks of stage 9

Copyright. Newspaper articles are not ours to republish. This notebook is built so that nothing but the first 15 words of an article ever reaches a file in output/: the map shows those words in the hover box and lets visitors search them, and nothing more. The full texts stay in data/, which is excluded from Git automatically (see stage 1).

1. Setup¶

Imports, folders and settings. Every setting you may want to change sits in this section. Files are read from and written to data/ and output/ next to the notebook.

In [ ]:
import base64
import hashlib
import html
import json
import os
import re
import textwrap
from pathlib import Path

import numpy as np
import pandas as pd
import plotly.graph_objects as go
from impresso import connect
from ollama import Client
from plotly.colors import sample_colorscale
from plotly.io import to_json
from plotly.offline import get_plotlyjs, get_plotlyjs_version
from sklearn.cluster import HDBSCAN
from sklearn.metrics import adjusted_rand_score
from sklearn.preprocessing import normalize
from tqdm import tqdm
from umap import UMAP
In [ ]:
# ── Folders, relative to this notebook ───────────────────────────────────────
DATA_DIR = Path("data")        # your corpus, with the full article text: never publish this folder
OUTPUT_DIR = Path("output")    # everything written here is safe to publish
DATA_DIR.mkdir(exist_ok=True)
OUTPUT_DIR.mkdir(exist_ok=True)
(DATA_DIR / ".gitignore").write_text("*\n")      # keeps data/ out of Git, whatever else you commit

INPUT_PATH = DATA_DIR / "articles.pkl"                 # your corpus
EMBEDDED_PATH = DATA_DIR / "articles_embedded.pkl"     # the same, plus one embedding per article
COORDS_PATH = OUTPUT_DIR / "umap_coordinates.csv"      # ids and map positions only, no text
LABELS_PATH = OUTPUT_DIR / "cluster_labels.json"       # the names given to the clusters
HTML_PATH = OUTPUT_DIR / "umap-labelled-map.html"      # the interactive map

# ── Columns of your DataFrame ────────────────────────────────────────────────
ID_COL, TEXT_COL, TITLE_COL = "id", "text.content", "text.title"
DATE_COL, PAPER_COL, LANG_COL = "meta.date", "meta.mediaId", "text.langCode"

# ── What is shown on the map ─────────────────────────────────────────────────
NEWSPAPERS = {                     # newspaper code -> name shown to visitors
    "JDG": "Journal de Genève",
    "GDL": "Gazette de Lausanne",
    "DBB": "Der Bund",
    "NZZ": "Neue Zürcher Zeitung",
}
YEAR_RANGE = (1940, 1999)          # pins the colour scale, so colours are comparable across runs
COLORSCALE = "Turbo"               # colour of the dots by year: blue (oldest) to red (most recent); try "Jet" or "Rainbow"
SNIPPET_WORDS = 15                 # how much of each article is shown: the publishable maximum

# ── Embedding service ────────────────────────────────────────────────────────
IMPRESSO_API_URL = "https://impresso-project.ch/public-api/v1"
EMBEDDING_COL = "embedding"            # what the service returns, as text
EMBEDDING_DEC = "embedding_impresso"   # the same, decoded into a numeric vector

Design¶

The fonts and layout of the maps follow the design of the website. The year is encoded by the continuous Turbo colour scale (COLORSCALE), so the oldest articles are blue and the most recent ones red. The same colour is used for the dots and for the box that appears when you hover over one.

In [3]:
INK, PAPER, GRID = "#333333", "#ffffff", "#e3ebe9"
FONT = "Google Sans Flex, Helvetica Neue, Arial, sans-serif"
MONO = "IBM Plex Mono, Menlo, Consolas, monospace"


def _hex(rgb):
    """'rgb(12, 34, 56)' (what plotly returns) -> '#0c2238'."""
    return "#%02x%02x%02x" % tuple(int(float(v)) for v in re.findall(r"[\d.]+", rgb))


def _text_on(colour):
    """Black or white text, whichever is more legible on this background colour."""
    r, g, b = (int(colour[i:i + 2], 16) for i in (1, 3, 5))
    return INK if 0.299 * r + 0.587 * g + 0.114 * b > 150 else "#ffffff"


Y0, Y1 = YEAR_RANGE
YEAR_STYLE = {}                    # year -> [background, text colour], shared by dots and hover boxes
for _year in range(Y0, Y1 + 1):
    _bg = _hex(sample_colorscale(COLORSCALE, [(_year - Y0) / (Y1 - Y0)])[0])
    YEAR_STYLE[_year] = [_bg, _text_on(_bg)]

Helper functions¶

Small functions used further down: decoding the embeddings and reducing them with UMAP (the first five), then turning the table into what the map shows (dates, newspaper names and the 15-word openings).

In [4]:
def decode_impresso_embedding(value):
    """Decode the service's '<model tag>:<base64 text of float32 numbers>' format into an array."""
    if not isinstance(value, str) or ":" not in value:
        return None
    _, payload = value.split(":", 1)
    try:
        return np.frombuffer(base64.b64decode(payload), dtype=np.float32).copy()
    except Exception:
        return None


def decode_impresso_column(df, raw_col="embedding", embedding_col="embedding_impresso"):
    """Decode a whole column of embeddings."""
    df[embedding_col] = df[raw_col].apply(decode_impresso_embedding)
    return df


def stack_embeddings(df, embedding_col):
    """Stack a column of vectors into one 2-D array. Rows without a valid vector are dropped.
    Returns (X, df_valid): the rows of X and df_valid line up."""
    valid = df[embedding_col].apply(lambda v: isinstance(v, (list, np.ndarray)))
    if (~valid).sum():
        print(f"  Dropping {(~valid).sum()} rows with a missing or invalid '{embedding_col}'")
    df_valid = df.loc[valid].copy()
    lengths = df_valid[embedding_col].apply(len)
    if lengths.nunique() > 1:
        raise ValueError(f"Embeddings of different lengths in '{embedding_col}': {lengths.value_counts().to_dict()}")
    return np.vstack(df_valid[embedding_col].to_list()).astype(np.float32), df_valid


def check_embeddings(df, embedding_cols):
    """Print the shape of each embedding column and how many rows had to be dropped."""
    for col in embedding_cols:
        X, df_valid = stack_embeddings(df, col)
        print(f"{col}: shape {X.shape}, {len(df) - len(df_valid)} dropped")


def run_umap(df, embedding_col, n_neighbors=15, min_dist=0.1, metric="cosine",
             normalize_vectors=True, random_state=42):
    """Fit a 2-D UMAP on one embedding column. Returns the valid rows with umap_x and umap_y added."""
    X, df_valid = stack_embeddings(df, embedding_col)
    if normalize_vectors:
        X = normalize(X, norm="l2")
    reducer = UMAP(n_neighbors=n_neighbors, min_dist=min_dist, metric=metric,
                   n_components=2, random_state=random_state)
    coords = reducer.fit_transform(X)
    df_valid = df_valid.copy()
    df_valid["umap_x"], df_valid["umap_y"] = coords[:, 0], coords[:, 1]
    return df_valid


def parse_dates(series):
    """Publication dates as timestamps (NaT when unreadable). Works on ISO strings, datetimes, with or without timezone."""
    return pd.to_datetime(series.astype(str).str[:10], errors="coerce")


def opening_words(text, n=SNIPPET_WORDS):
    """The first n words of an article: the only part of the text that is ever published."""
    words = text.split() if isinstance(text, str) else []
    return " ".join(words[:n]) + (" …" if len(words) > n else "")

2. Embed¶

An embedding turns a text into a list of numbers (here, 256 of them) so that texts with similar meanings get similar numbers. Everything that follows (the map, the clusters) is built on these vectors.

The vectors come from the embedding model of the Impresso project (m-GTE), through its public API. The call is slow for a large corpus, so the loop saves its progress regularly: if it is interrupted, run the cell again and it resumes where it stopped.

In [5]:
impresso_session = connect(IMPRESSO_API_URL)    # asks you to log in the first time
Click on the following link to access the login page: https://dev.impresso-project.ch/datalab/token
 - 🔤 Enter your email/password on this page.
 - 🔑 Once logged in, a secret token will be generated for you.
 - 📋 Copy this token and paste it into the input field below. Then press "Enter". 👇🏼.

🎉 You are now connected to the Impresso API!  🎉
🔗 Using API: https://dev.impresso-project.ch/public-api/v1
In [8]:
def embed_text(text, target="text"):
    """Embedding of one text, or None if the service fails (the article is retried on the next run)."""
    try:
        return impresso_session.tools.embed_text(text, target)
    except Exception:
        return None


df = pd.read_pickle(EMBEDDED_PATH if EMBEDDED_PATH.exists() else INPUT_PATH).reset_index(drop=True)
if EMBEDDING_COL not in df.columns:
    df[EMBEDDING_COL] = None

todo = df.index[df[EMBEDDING_COL].isna()]
print(f"{len(df):,} articles, {len(todo):,} still to embed")

for n, i in enumerate(tqdm(todo, desc="embedding"), 1):
    df.at[i, EMBEDDING_COL] = embed_text(df.at[i, TEXT_COL])
    if n % 1000 == 0:
        df.to_pickle(EMBEDDED_PATH)             # checkpoint
df.to_pickle(EMBEDDED_PATH)
6,154 articles, 1 still to embed
embedding: 100%|██████████| 1/1 [00:00<00:00,  1.92it/s]

The service returns each vector as text. We decode it into numbers, and drop the articles for which no embedding could be obtained (they cannot be placed on the map).

In [9]:
if EMBEDDING_DEC not in df.columns:
    df = decode_impresso_column(df, raw_col=EMBEDDING_COL, embedding_col=EMBEDDING_DEC)

df = df[df[EMBEDDING_DEC].notna()].reset_index(drop=True)
if df.empty:
    raise ValueError("No valid embedding left: check that the embedding service answered (stage 2).")

print(f"{len(df):,} articles with a valid embedding")
print("vector length:", len(df[EMBEDDING_DEC].iloc[0]))
6,153 articles with a valid embedding
vector length: 256

3. Reduce to two dimensions with UMAP¶

A vector of 256 numbers cannot be drawn. UMAP projects the vectors onto a plane while trying to keep neighbours together: articles whose vectors are close stay close on the map. The axes themselves have no meaning; only distances between dots do.

Two settings shape the result:

  • N_NEIGHBORS: how many neighbours each article looks at. Low values bring out fine local detail; high values keep more of the corpus's overall structure.
  • MIN_DIST: how tightly dots may be packed. Low values give compact clumps; high values spread them out.

We use the cosine distance, the usual choice for text embeddings, on vectors normalised to length 1. The random seed is fixed, so the same input gives the same map.

In [10]:
N_NEIGHBORS = 50
MIN_DIST = 0.01
UMAP_METRIC = "cosine"
NORMALIZE_VECTORS = True     # recommended with the cosine metric

check_embeddings(df, [EMBEDDING_DEC])

coords = run_umap(
    df,
    embedding_col=EMBEDDING_DEC,
    n_neighbors=N_NEIGHBORS,
    min_dist=MIN_DIST,
    metric=UMAP_METRIC,
    normalize_vectors=NORMALIZE_VECTORS,
)

coords[[ID_COL, "umap_x", "umap_y"]].to_csv(COORDS_PATH, index=False)   # positions only, no text
embedding_impresso: shape (6153, 256), 0 dropped
/opt/anaconda3/envs/py314-phd/lib/python3.14/site-packages/umap/umap_.py:1952: UserWarning: n_jobs value 1 overridden to 1 by setting random_state. Use no seed for parallelism.
  warn(
OMP: Info #276: omp_set_nested routine deprecated, please use omp_set_max_active_levels instead.

4. First look: the map, coloured by year¶

Two functions turn the table into a figure. They are used for every map in this notebook, so the dots and the hover boxes look the same everywhere.

  • prepare_points keeps what a visitor will see about each article: the newspaper, the date, the cluster name (once we have one) and the first 15 words. It does not carry the rest of the text any further.
  • make_figure draws the map. Hovering over a dot shows a box coloured like the dot, with the cluster, the newspaper, the date and the opening of the article.

Articles whose date cannot be read are left out of the maps.

In [11]:
def prepare_points(d, cluster_names=None):
    """What a visitor may see about each article. `d` needs umap_x, umap_y and, once clustered, `cluster`.
    `cluster_names`: a column or series with the name of each article's cluster (optional)."""
    dates = parse_dates(d[DATE_COL])
    keep = dates.notna().to_numpy()
    if not keep.all():
        print(f"{(~keep).sum()} article(s) with an unreadable date left out of the map")
    d, dates = d.loc[keep], dates[keep]

    pts = pd.DataFrame({
        "x": d["umap_x"].to_numpy(),
        "y": d["umap_y"].to_numpy(),
        "year": dates.dt.year.to_numpy(),
        "date": dates.dt.strftime("%Y-%m-%d").to_numpy(),
        "date_label": [f"{t.day} {t:%B %Y}" for t in dates],
        "paper": d[PAPER_COL].map(NEWSPAPERS).fillna(d[PAPER_COL].astype(str)).to_numpy(),
        "snippet": [opening_words(t) for t in d[TEXT_COL]],
        "cluster": d["cluster"].to_numpy() if "cluster" in d else -1,
        "cluster_name": (np.asarray(cluster_names)[keep] if cluster_names is not None
                         else np.full(len(d), "", dtype=object)),
    }, index=d.index)

    # the text of the native hover box, with the opening wrapped to a readable width
    wrapped = [html.escape(textwrap.fill(s, 44)).replace("\n", "<br>") for s in pts["snippet"]]
    pts["hover"] = [
        (f"<b>{html.escape(str(c))}</b><br>" if c else "") + f"{p} · {t}<br><i>{w}</i>"
        for c, p, t, w in zip(pts["cluster_name"], pts["paper"], pts["date_label"], wrapped)
    ]
    return pts


def cluster_label_anchors(pts, names):
    """One label per cluster: its name, sat on the article nearest the cluster's centre so it always sits on data."""
    sizes = pts["cluster"].value_counts().drop(-1, errors="ignore")
    anchors = []
    for c, name in names.items():
        sub = pts[pts["cluster"] == c]
        if sub.empty:
            continue
        cx, cy = sub["x"].median(), sub["y"].median()
        nearest = sub.iloc[np.argmin(np.hypot(sub["x"] - cx, sub["y"] - cy))]
        size = float(np.interp(len(sub), [sizes.min(), sizes.max()], [10, 16]))
        anchors.append((nearest["x"], nearest["y"], name, size))
    return anchors


def _ring(radius, n=8):
    """n pixel offsets evenly spaced on a circle."""
    step = 2 * np.pi / n
    return [(round(radius * np.cos(k * step), 2), round(radius * np.sin(k * step), 2)) for k in range(n)]


GLOW_LAYERS = [(1, 1.0), (2, 0.7), (3.5, 0.35)]   # (radius in px, white opacity) of the halo behind each label


def add_label(fig, x, y, text, size):
    """A label with a white halo. Plotly has no text outline, so the halo is the same text in white,
    repeated on a ring of small offsets and stacked under the dark text."""
    text = f"<b>{html.escape(text)}</b>"
    for radius, opacity in GLOW_LAYERS:
        for dx, dy in _ring(radius * size / 13):
            fig.add_annotation(x=x, y=y, xshift=dx, yshift=dy, showarrow=False, text=text,
                               font=dict(size=size, color=f"rgba(255,255,255,{opacity})", family=FONT))
    fig.add_annotation(x=x, y=y, showarrow=False, text=text, font=dict(size=size, color=INK, family=FONT))


def corpus_caption(n):
    papers = ", ".join(f"<i>{p}</i>" for p in NEWSPAPERS.values())
    return f"Articles published in {papers}, with <b>n={n:,}</b>. UMAP parameters: n_neighbors={N_NEIGHBORS}, min_dist={MIN_DIST}."


def make_figure(pts, anchors=(), title=None, size=4, opacity=0.7, interactive_export=False):
    """The map. With interactive_export=True the figure is built for the standalone page of stage 10:
    it fills its container, has no native hover box (the page draws its own) and has an empty second
    trace that the page uses to highlight search results."""
    decades = list(range(Y0 // 10 * 10, Y1 + 1, 10))
    dots = dict(
        size=size, opacity=opacity, color=pts["year"], colorscale=COLORSCALE, cmin=Y0, cmax=Y1,
        colorbar=dict(title=dict(text="decade", font=dict(family=MONO, size=12)), thickness=12, len=0.6,
                      tickvals=[d + 5 for d in decades], ticktext=[f"{d}s" for d in decades],
                      outlinecolor=INK, outlinewidth=1.5, tickfont=dict(family=MONO, size=11)),
    )
    fig = go.Figure()
    if interactive_export:
        fig.add_trace(go.Scattergl(x=pts["x"], y=pts["y"], mode="markers", marker=dots, hoverinfo="none"))
        highlight = dict(dots, size=size + 1, opacity=1, colorbar=None, showscale=False, color=[])
        fig.add_trace(go.Scattergl(x=[], y=[], mode="markers", marker=highlight, hoverinfo="none"))
    else:
        fig.add_trace(go.Scattergl(
            x=pts["x"], y=pts["y"], mode="markers", marker=dots,
            text=pts["hover"], hovertemplate="%{text}<extra></extra>",
            hoverlabel=dict(                      # the hover box takes the colour of its dot
                bgcolor=[YEAR_STYLE[min(max(y, Y0), Y1)][0] for y in pts["year"]],
                font=dict(family=FONT, size=12, color=[YEAR_STYLE[min(max(y, Y0), Y1)][1] for y in pts["year"]]),
                bordercolor=INK, align="left",
            ),
        ))

    for x, y, name, label_size in anchors:
        add_label(fig, x, y, name, label_size)

    axis = dict(showgrid=True, gridcolor=GRID, zeroline=True, zerolinecolor=GRID, showline=True,
                linecolor=INK, linewidth=1.5, ticks="outside", tickcolor=INK, tickfont=dict(size=10))
    fig.update_layout(
        template="none", paper_bgcolor=PAPER, plot_bgcolor=PAPER, showlegend=False,
        font=dict(family=FONT, color=INK, size=12),
        xaxis=dict(title="UMAP 1", **axis), yaxis=dict(title="UMAP 2", **axis),
    )
    if interactive_export:
        fig.update_layout(autosize=True, margin=dict(l=48, r=12, t=12, b=44), dragmode="pan")
    else:
        fig.update_layout(
            width=1100, height=750, margin=dict(l=60, r=20, t=90, b=50),
            title=dict(text=title, x=0.01, y=0.97, font=dict(family=MONO, size=18)),
            annotations=list(fig.layout.annotations) + [dict(
                text=corpus_caption(len(pts)), xref="paper", yref="paper", x=0, xanchor="left",
                y=1.02, yanchor="bottom", showarrow=False, align="left", font=dict(size=10, color="#666666"))],
        )
    return fig
In [12]:
pts = prepare_points(coords)
fig_year = make_figure(pts, title="UMAP of the corpus, coloured by year")
fig_year.show()

5. Cluster the map with HDBSCAN¶

The map shows groups of articles. To name them, we first need to say precisely which articles belong to which group. HDBSCAN finds dense regions of points and labels each one as a cluster. Points in no dense region are left out as noise (cluster -1) and never get a name.

By default the clustering runs on the 2-D coordinates you see, so the groups you read off the map are exactly the groups that get named. The alternative, "umap10", clusters a separate 10-dimensional UMAP: more faithful to the embeddings, but a cluster may then look scattered on the 2-D map.

The settings that matter:

  • MIN_CLUSTER_SIZE: the smallest group that counts as a cluster. Lower values give more, finer clusters.
  • MIN_SAMPLES: higher values are more conservative and leave more points unclustered.
  • SELECTION: "eom" favours fewer, larger clusters, "leaf" gives many fine ones.
In [13]:
CLUSTER_SPACE = "umap2d"          # "umap2d" or "umap10", see above
MIN_CLUSTER_SIZE = max(25, len(coords) // 100)
MIN_SAMPLES = 50
SELECTION = "eom"

dfc = coords.copy().reset_index(drop=True)

if CLUSTER_SPACE == "umap2d":
    Z = dfc[["umap_x", "umap_y"]].to_numpy()
else:
    X, _ = stack_embeddings(dfc, EMBEDDING_DEC)    # same rows, same order as dfc
    Z = UMAP(n_components=10, n_neighbors=N_NEIGHBORS, min_dist=0.0, metric=UMAP_METRIC,
             random_state=42).fit_transform(normalize(X))

hdb = HDBSCAN(min_cluster_size=MIN_CLUSTER_SIZE, min_samples=MIN_SAMPLES, cluster_selection_method=SELECTION, copy=True)
dfc["cluster"] = hdb.fit_predict(Z)          # -1 = noise
dfc["cluster_prob"] = hdb.probabilities_     # how firmly each article belongs to its cluster

ids = sorted(c for c in dfc["cluster"].unique() if c != -1)
assert ids, "HDBSCAN found no cluster: lower MIN_CLUSTER_SIZE or MIN_SAMPLES"
print(f"{len(ids)} clusters, {(dfc['cluster'] == -1).mean():.1%} of articles unclustered")
print(dfc["cluster"].value_counts().drop(-1, errors="ignore").describe().round(0).to_string())
20 clusters, 41.6% of articles unclustered
count     20.0
mean     180.0
std      137.0
min       64.0
25%       95.0
50%      128.0
75%      221.0
max      643.0

A quick look at the raw clusters before spending any language-model calls on them (grey: unclustered).

In [14]:
def cluster_colours(cluster_ids):
    """A distinct colour per cluster, taken from the colour scale at golden-ratio steps so neighbours differ."""
    return {c: _hex(sample_colorscale(COLORSCALE, [(k * 0.618034) % 1])[0]) for k, c in enumerate(cluster_ids)}


palette = cluster_colours(ids)
fig_cl = go.Figure(go.Scattergl(
    x=dfc["umap_x"], y=dfc["umap_y"], mode="markers",
    marker=dict(size=3, opacity=0.6, color=[palette.get(c, "#cccccc") for c in dfc["cluster"]]),
    text=dfc["cluster"].astype(str), hovertemplate="cluster %{text}<extra></extra>",
))
fig_cl.update_layout(
    template="none", width=900, height=600, paper_bgcolor=PAPER, plot_bgcolor=PAPER,
    font=dict(family=FONT, color=INK, size=12), margin=dict(l=60, r=20, t=70, b=50),
    title=dict(text=f"HDBSCAN clusters (min_cluster_size={MIN_CLUSTER_SIZE}, min_samples={MIN_SAMPLES})",
               x=0.01, font=dict(family=MONO, size=16)),
    xaxis=dict(title="UMAP 1", gridcolor=GRID, linecolor=INK), yaxis=dict(title="UMAP 2", gridcolor=GRID, linecolor=INK),
)
fig_cl.show()

6. Pick representative articles¶

A language model cannot read thousands of articles per cluster, so it gets N_ARTICLES_PER_CLUSTER of them: those that belong most firmly to the cluster (strongest HDBSCAN membership) come first, and ties are broken by distance to the cluster's centre. These are the most typical articles of each group.

In [15]:
N_ARTICLES_PER_CLUSTER = 30

dfc["d_centre"] = np.nan
for c in ids:
    m = (dfc["cluster"] == c).to_numpy()
    dfc.loc[m, "d_centre"] = np.linalg.norm(Z[m] - np.median(Z[m], axis=0), axis=1)

reps = {
    c: dfc[dfc["cluster"] == c]
        .sort_values(["cluster_prob", "d_centre"], ascending=[False, True])
        .head(N_ARTICLES_PER_CLUSTER)
    for c in ids
}

7. Name the clusters with a language model¶

For each cluster, the model reads its representative articles and answers with a short label (2–4 words). The instructions it receives are in LABEL_SYSTEM_PROMPT; adapt them to your corpus.

  • One call per cluster, at temperature 0 so the answer is as repeatable as possible.
  • Labels are cached in output/cluster_labels.json, under a hash of the model, the prompt and the articles. Re-running the cell only queries clusters whose inputs changed, and the published labels stay reproducible.
  • What leaves your machine. Each call sends the title and the first ARTICLE_CHAR_LIMIT characters of the representative articles to the model's server. Make sure your data licence allows this, or run a local model instead (OLLAMA_HOST = "http://localhost:11434").
  • Access. Set an OLLAMA_API_KEY environment variable for Ollama Cloud. No key is needed for a local model.
In [ ]:
LABELLING_MODEL = "glm-5.3-flash:cloud"
LABELLING_THINK = False                 # True if the model supports thinking and you want it
ARTICLE_CHAR_LIMIT = 1000               # characters of each article sent to the model
OLLAMA_HOST = "https://ollama.com"

LABEL_SYSTEM_PROMPT = """
You are an expert historian of Swiss banking and finance (1940–1999).
Given a numbered list of newspaper articles in French or German (may contain OCR errors) that all belong to one cluster of articles mentioning the Swiss Bankers Association:
Return ONLY: { "label": "your topic label" }
RULES: 2–4 words, lower-case English noun phrase, monograph-section specificity.
Do not use "banking" (implicit) or "swiss" unless genuinely distinctive.
""".strip()

api_key = os.environ.get("OLLAMA_API_KEY")
ollama_client = Client(host=OLLAMA_HOST, headers={"Authorization": f"Bearer {api_key}"} if api_key else None)

cache = json.loads(LABELS_PATH.read_text()) if LABELS_PATH.exists() else {}

labels = {}
for c in tqdm(ids, desc="labelling clusters"):
    user_content = "Articles:\n\n" + "\n\n".join(
        f"[{k}] ({r[PAPER_COL]}, {str(r[DATE_COL])[:4]}) {r[TITLE_COL]}\n"
        f"{' '.join(str(r[TEXT_COL]).split())[:ARTICLE_CHAR_LIMIT]}"
        for k, (_, r) in enumerate(reps[c].iterrows(), 1)
    )
    key = f"{LABELLING_MODEL}:{hashlib.sha1((LABEL_SYSTEM_PROMPT + user_content).encode()).hexdigest()}"

    if key not in cache:
        for attempt in range(3):
            try:
                resp = ollama_client.chat(
                    LABELLING_MODEL,
                    messages=[
                        {"role": "system", "content": LABEL_SYSTEM_PROMPT},
                        {"role": "user", "content": user_content},
                    ],
                    think=LABELLING_THINK,
                    options={"temperature": 0.0},
                )
                raw = re.sub(r"<think>.*?</think>", "", resp.message.content, flags=re.DOTALL)
                cache[key] = json.loads(re.search(r"\{.*\}", raw, re.S).group(0))["label"].strip()
                LABELS_PATH.write_text(json.dumps(cache, ensure_ascii=False, indent=1))
                break
            except Exception as e:
                print(f"    [cluster {c}, attempt {attempt + 1}/3] {e}")
    labels[c] = cache.get(key, "unlabelled")

dfc["cluster_label"] = dfc["cluster"].map(labels).fillna("unclustered")

cluster_summary = pd.DataFrame({
    "cluster": ids,
    "n": [(dfc["cluster"] == c).sum() for c in ids],
    "label": [labels[c] for c in ids],
})
dupes = cluster_summary.loc[cluster_summary["label"].duplicated(keep=False), "label"].unique()
if len(dupes):
    print("Same label given to several clusters (merge them, or tighten the prompt):", list(dupes))
cluster_summary
labelling clusters:   0%|          | 0/20 [00:00<?, ?it/s]
    [cluster 0, attempt 1/3] Extra data: line 1 column 38 (char 37)
labelling clusters:  20%|██        | 4/20 [00:24<01:38,  6.16s/it]
    [cluster 4, attempt 1/3] Extra data: line 1 column 38 (char 37)
    [cluster 4, attempt 2/3] Extra data: line 1 column 32 (char 31)
labelling clusters:  30%|███       | 6/20 [01:05<02:57, 12.65s/it]
    [cluster 6, attempt 1/3] Extra data: line 1 column 51 (char 50)
labelling clusters:  50%|█████     | 10/20 [01:29<01:07,  6.78s/it]
    [cluster 10, attempt 1/3] Extra data: line 1 column 47 (char 46)
labelling clusters:  75%|███████▌  | 15/20 [01:50<00:23,  4.64s/it]
    [cluster 15, attempt 1/3] Extra data: line 1 column 31 (char 30)
labelling clusters: 100%|██████████| 20/20 [02:32<00:00,  7.61s/it]
Out[ ]:
cluster n label
0 0 76 cashless payment systems
1 1 181 us mutual legal assistance treaty
2 2 64 sevillana de electricidad dividends
3 3 102 american dividend payments, 1943–1953
4 4 124 dividend and bond notices
5 5 84 bankers association staff recruitment
6 6 281 holocaust-era dormant accounts controversy
7 7 223 swiss bankers association leadership
8 8 377 eastern european debt settlements
9 9 118 private discount rate adjustments
10 10 91 anti-money laundering legislation
11 11 139 due diligence agreement (VSB)
12 12 192 voluntary credit limitation agreements
13 13 67 registered shares and takeover defense
14 14 96 european integration policy
15 15 220 banking law revision
16 16 643 federal finance reform debates
17 17 127 mortgage interest rate policy
18 18 260 bankers' day annual meeting
19 19 128 banks as employers

8. The labelled map¶

The same map as in stage 4, now with the clusters' names written on it. Each label sits on the article closest to its cluster's centre, and its size grows with the size of the cluster. Hover over a dot to see which cluster it belongs to.

In [22]:
pts = prepare_points(dfc, cluster_names=dfc["cluster_label"])
anchors = cluster_label_anchors(pts, labels)

fig_lab = make_figure(pts, anchors, title="UMAP of SBA coverage, 1940–1999: clusters named by an LLM")
fig_lab.show()

9. Check the labels¶

The labels are a reading aid, not a finding. Before relying on them, three checks:

  1. Who is in each cluster? A cluster made 95% of one newspaper may reflect a house style or an OCR quirk rather than a topic. With two German and two French titles, language splits show up here too.
  2. Skim a few articles per cluster. Does the label hold?
  3. Do the clusters survive other settings? If small changes to HDBSCAN reshuffle everything, the clusters are fragile.
In [23]:
PAPER_SHARE_FLAG = 0.9    # flag clusters where one newspaper makes up more than this share ...
DECADE_SHARE_FLAG = 0.7   # ... or one decade does (legitimate for event-driven topics, but worth a look)

dfc["year"] = parse_dates(dfc[DATE_COL]).dt.year

rows = []
for c in ids:
    sub = dfc[dfc["cluster"] == c]
    paper = sub[PAPER_COL].value_counts(normalize=True)
    decade = (sub["year"] // 10 * 10).value_counts(normalize=True)
    rows.append({
        "cluster": c, "label": labels[c], "n": len(sub),
        "top_paper": paper.index[0], "top_paper_share": paper.iloc[0],
        "top_decade": int(decade.index[0]), "top_decade_share": decade.iloc[0],
        "median_year": sub["year"].median(), "share_de": (sub[LANG_COL] == "de").mean(),
        "mean_membership": sub["cluster_prob"].mean(),
    })
review = pd.DataFrame(rows)
review["flag"] = np.where(
    (review["top_paper_share"] > PAPER_SHARE_FLAG) | (review["top_decade_share"] > DECADE_SHARE_FLAG), "check", ""
)
review.round(2)
Out[23]:
cluster label n top_paper top_paper_share top_decade top_decade_share median_year share_de mean_membership flag
0 0 cashless payment systems 76 JDG 0.28 1980 0.59 1987.0 0.50 0.95
1 1 us mutual legal assistance treaty 181 NZZ 0.39 1980 0.61 1983.0 0.54 0.73
2 2 sevillana de electricidad dividends 64 JDG 0.80 1980 0.34 1977.5 0.06 0.99
3 3 american dividend payments, 1943–1953 102 JDG 0.69 1940 0.89 1945.0 0.30 0.91 check
4 4 dividend and bond notices 124 NZZ 0.81 1940 0.46 1950.0 0.81 0.87
5 5 bankers association staff recruitment 84 NZZ 0.58 1970 0.37 1978.0 0.73 0.97
6 6 holocaust-era dormant accounts controversy 281 NZZ 0.50 1990 0.95 1997.0 0.52 0.82 check
7 7 swiss bankers association leadership 223 NZZ 0.36 1980 0.33 1978.0 0.62 0.88
8 8 eastern european debt settlements 377 JDG 0.34 1940 0.48 1950.0 0.49 0.79
9 9 private discount rate adjustments 118 NZZ 0.31 1980 0.49 1980.0 0.50 0.92
10 10 anti-money laundering legislation 91 JDG 0.34 1990 0.55 1990.0 0.55 0.96
11 11 due diligence agreement (VSB) 139 NZZ 0.28 1980 0.58 1984.0 0.47 0.89
12 12 voluntary credit limitation agreements 192 JDG 0.34 1970 0.48 1975.0 0.42 0.84
13 13 registered shares and takeover defense 67 NZZ 0.52 1980 0.42 1988.0 0.60 1.00
14 14 european integration policy 96 JDG 0.50 1990 0.78 1992.0 0.29 0.89 check
15 15 banking law revision 220 NZZ 0.32 1980 0.57 1984.0 0.50 0.92
16 16 federal finance reform debates 643 NZZ 0.36 1970 0.29 1973.0 0.54 0.63
17 17 mortgage interest rate policy 127 NZZ 0.32 1980 0.39 1980.0 0.54 0.90
18 18 bankers' day annual meeting 260 NZZ 0.36 1970 0.22 1968.0 0.56 0.98
19 19 banks as employers 128 JDG 0.35 1980 0.45 1986.0 0.47 0.97

Skimming a few random articles per cluster. Only the newspaper, the date and the opening words are printed, never the full text.

In [24]:
N_SKIM = 5
for c in ids:
    sub = dfc[dfc["cluster"] == c]
    print(f"\n[{c}] {labels[c]}  (n={len(sub)})")
    for _, r in sub.sample(min(N_SKIM, len(sub)), random_state=42).iterrows():
        print(f"   {NEWSPAPERS.get(r[PAPER_COL], r[PAPER_COL])}, {str(r[DATE_COL])[:10]} | {opening_words(r[TEXT_COL])}")
[0] cashless payment systems  (n=76)
   Der Bund, 1987-05-26 | Bancomat-Code nicht knackbar bee. Das Bancomat-System ist nach den Worten der Schweizerischen Bankiervereinigung «technisch sicher». …
   Der Bund, 1988-02-26 | PTT und Migros starten Pilotversuch im Shoppyland Ab Montag können Inhaber eines Postscheckkontos bargeldlos einkaufen-Bonitätslimite …
   Der Bund, 1987-05-21 | Bundesgerichtsurteil stellt bisherige Schadenregelung in Frage Wie sicher sind Bancomaten? ap. Die Schweizer Banken lassen …
   Der Bund, 1986-07-08 | Risiken meiden Nahezu die Hälfte der abhandenkommenden Zahlungsmittel wie Euroche-ques Kreditkarten oder Travellers Che-ques sei …
   Gazette de Lausanne, 1980-03-24 | Nouveaux chèques de voyage européens INTERVIEW DU PRÉSIDENT DU COMITÉ DES « SWISS BANKERS TRAVELLERS …

[1] us mutual legal assistance treaty  (n=181)
   Der Bund, 1983-11-17 | Bereits eine Bankenvereinbarung Diese Tatsache gab im letzten Jahr An-lass zur Schaffung der «Konvention XVI …
   Neue Zürcher Zeitung, 1986-02-04 | Umstrittene Insidervorlage Das Interesse der Banken an der Lösung des Insiderproblems Von Jean-Paul Chapuis, Delegiertem …
   Journal de Genève, 1982-01-09 | Les banques suisses condamnent l'usage d'informations privilégiées en bourse Les banques suisses ne veulent pas …
   Der Bund, 1952-06-19 | Vermischte Anfragen im Nationalrat Sitzung von Mittwochnachmittag, 18. Juni. — Vorsitz: Präsident Renold (Aargau) -i-Zu …
   Journal de Genève, 1988-07-01 | BOURSE Nouveau scandale d'initiés aux USA : ramifications possibles en Suisse La Securities and Exchange …

[2] sevillana de electricidad dividends  (n=64)
   Journal de Genève, 1984-06-15 | AFFAIRES IMMOBILIERES En raison de son expansion "constante bureaux de haut standing de 2000 à …
   Journal de Genève, 1980-12-17 | COMPANIA SEVILLANA DE fLECTRICIDAD ACH ~ TE vIeil or, argent, vIeux biJoux COMPTOIR DESM ~ …
   Neue Zürcher Zeitung, 1964-05-20 | Compania Sevillana de Electricidad Dividendenzahlung Den Aktionären wird hiemit bekanntgegeben, daß die Generalversammlung der oben …
   Journal de Genève, 1968-12-17 | DAGGAFONTEIN MINES, LIMITED (enregistrée en République sud-africaine) Paiement du coupon de dividende No 72 Les …
   Gazette de Lausanne, 1982-06-10 | AVIS FINANCIERS COMPANIA SEVILLANA DE ELECTRICIDAD (Numéro de valeur 467.093) Paiement de dividende Le Conseil …

[3] american dividend payments, 1943–1953  (n=102)
   Neue Zürcher Zeitung, 1947-05-12 | für die Kotierung an den Boörsen von Zürich, Basel und Genf von hfl. 100 724 …
   Journal de Genève, 1947-11-01 | Avis Financiers La Rente Immobilière Le dividende de l'exercice 1946 / 47, au montant de …
   Journal de Genève, 1945-09-20 | General Motors Corporation Common shares •• Paiement de dividende Avis est donné aux porteurs des …
   Journal de Genève, 1945-11-02 | Avis Financiers General Electric Company Common shares / Paiement de dividende Avis est donné aux …
   Journal de Genève, 1943-10-21 | Montgomery Ward & Co, Chicago Common shares / Paiement de dividende Avis est donné aux …

[4] dividend and bond notices  (n=124)
   Neue Zürcher Zeitung, 1978-10-19 | c Management by Fein-Kaller ä ■'7?^* W|l, JlM l^irFLWfV l,J^y‘?^Hg^^v>y ia L‘'‘"-> |,,''^'.7, I!’ A …
   Neue Zürcher Zeitung, 1967-11-22 | I S »*'•..XV-..-.-»'fäf •• ’M" • ■ ■ • 1»'•: ’r-''•'T, p fwMii oPlra-jySB k; …
   Neue Zürcher Zeitung, 1969-12-08 | BFOr jedermann HAMMOND-<fjF~ Orgel-Unterricht'in kleinen Gruppen Neue Abendkurse Eiffi ||Q (10 Wochen-Lektionen Fr. 60.—) Kursbeginn …
   Neue Zürcher Zeitung, 1972-11-03 | ‘■!»■■»—■ ■ i m | i. t i wn» i |, | n ■ ■ …
   Neue Zürcher Zeitung, 1944-08-02 | —— :; 2 S t 9 2.9 —— s c Ein italienischer Spitzenfilm —2 1 …

[5] bankers association staff recruitment  (n=84)
   Journal de Genève, 1978-06-03 | L'Association Suisse des Banquiers cherche pour son secrétariat, à Baie une secrétaire pour la correspondance …
   Der Bund, 1965-05-15 | Die SCHWEIZERISCHE BANKIERVEREINIGUNG sucht für ihr Sekretariat in Basel jungen juristischen volkswirtschaftlichen Mitarbeiter Müttersprache Deutsch, …
   Neue Zürcher Zeitung, 1999-11-20 | Fachjournalist BR <, I BANKVEREINIGUNG Mitte Dreissig, D/E/F/Sp,lic phU.,rmt langjäh-\ nger Erfahrung im Zeitschriften-(Fach-und Rnrha\ …
   Neue Zürcher Zeitung, 1970-01-28 | C.. S „„ Schweizerische Bankiervereinigung W JF i 1 H StO Is ®ffis 1 > …
   Neue Zürcher Zeitung, 1970-01-18 | durch Bedeutende Lebensversicherungsgesellschaft sucht ISB Hans Fehr AG 8305 Dietlikon/Zürich Generalagenten FEHR, ein Begriff auf …

[6] holocaust-era dormant accounts controversy  (n=281)
   Neue Zürcher Zeitung, 1996-09-13 | Nichts Neues in der Nazi-Gold-Affäre? Stellungnahme des Israelitischen Gemeindebundes Bern, 11. Sept, (sda) Der vom …
   Journal de Genève, 1996-10-19 | Le président de la Commission mixte sur les fonds juifs réfute les accusations du sénateur …
   Journal de Genève, 1996-10-19 | Paul Volcker répond à M. D'Amato sur les fonds juife Selon son président, la commission …
   Journal de Genève, 1997-07-07 | Affaire des fonds juifs : Flavio Cotti a poussé Carlo Jagmetti à partir L'ancien ambassadeur …
   Neue Zürcher Zeitung, 1996-05-24 | Programmhinweise DRS, 20.30 Quer Statistisch gesehen fahrt jeder Lokführer im Laufe seines Berufslebens mindestens einen …

[7] swiss bankers association leadership  (n=223)
   Neue Zürcher Zeitung, 1986-11-11 | Personalien Schweizerische Bankiervereinigung V Der Verwaltungsrat der Schweizerischen Bankiervereinigung hat Professor Niklaus Blattner zum Volkswirtschaftlichen …
   Der Bund, 1959-07-18 | Direktor Ernst Blaser 60jährig Heute, am 17. Juli, vollendet Herr Ernst Blaser, Mitglied der Zentraldirektion …
   Neue Zürcher Zeitung, 1976-12-04 | Hans Vontobel zum 60. Geburtstag Am 4. Dezember feiert Dr. Hans Vontobel seinen 60. Geburtstag. …
   Journal de Genève, 1991-02-07 | HOMMAGE Edmond Barbey Avec Edmond Barbey, récemment décédé à un âge avancé, disparaît une des …
   Journal de Genève, 1965-03-27 | Nouveau'président de l'Associàtion , suisse des banquier ~ , {Com.) M. C. de Loës, a …

[8] eastern european debt settlements  (n=377)
   Journal de Genève, 1951-05-29 | La suppression des affldavits A partir du 25 mai 1951, il n'est plus exigé d'affidavits …
   Journal de Genève, 1955-12-22 | Emprunts en francs suisses de débiteurs autrichiens Dans l'annexe I E de l'accord de Londres …
   Gazette de Lausanne, 1940-02-02 | Le marché des emprunts en Suisse en 1939 Les émissions et les conversions d'emprunt ont …
   Neue Zürcher Zeitung, 1954-09-22 | \J i r— U rrr J. BADEX W • BAORAGAZ • • Das Badhotel • …
   Der Bund, 1950-10-26 | AUFRUF betreffend Anmeldung schweizerischer Vermögenswerte, Rechte, Forderungen und Interessen In Rumänien Schweizerische und liechtensteinische Personen, …

[9] private discount rate adjustments  (n=118)
   Neue Zürcher Zeitung, 1982-09-02 | Inländische Finanzmärkte 4¾%-Anleihe der KB Schwyz (Reiner) Die Kantonalbank Schwyz legt vom 7. bis 13. …
   Journal de Genève, 1975-10-31 | Baisse des taux pour le financement des réserves obligatoires • Baie, 30. — (ATS) Tenant …
   Der Bund, 1989-04-28 | Neue Zinsrunde bei Kassenobligationen sda. Nach der Schweizerischen Bankgesellschaft ziehen auch der Schweizerische Bankverein (SBV) …
   Neue Zürcher Zeitung, 1982-10-16 | Fragwürdige Mischrechnung Höhere Depotgebühren für Wertschriften bb. Mit Wirkung ab 1. Januar 1983 haben Depotinhaber …
   Neue Zürcher Zeitung, 1982-02-02 | Höhere Bankspesen für Zahlungsaufträge Gemäss einer Konvention der Schweizerischen Bankiervereinigung wurden seit März 1978 einheitlich …

[10] anti-money laundering legislation  (n=91)
   Neue Zürcher Zeitung, 1990-09-04 | Ist die Schweiz das Paradies der Geldwäscher? Walter Frehner, Vorsitzender der Geschäftsleitung des Schweizerischen sind …
   Neue Zürcher Zeitung, 1976-07-14 | Warnung vor falschen DM-Reisechecks Die Schweizerische Bankiervereinigung teilt mit: Seit einiger Zeit werden gefälschte deutsche …
   Gazette de Lausanne, 1989-02-23 | L'ABS prête à de nouvelles mesures (AFP).-L'Association suisse des banquiers (ASB) s'est déclarée favorable à …
   Journal de Genève, 1988-01-22 | ; > I ~ J, ~, ~ ~. ~ SUISSE Blanchissage d'argent sate : vides …
   Neue Zürcher Zeitung, 1995-10-25 | Erstaunen bei der Bankiervereinigung Basel, 24. Okt. (ap) Die Schweizerische Bankiervereinigung hat den Bericht der …

[11] due diligence agreement (VSB)  (n=139)
   Journal de Genève, 1990-04-11 | BANQUES Clients des banques : vers l'abolition complète de l'anonymat en Suisse La Commission fédérale …
   Journal de Genève, 1977-08-24 | Convention avec la BNS : presque toutes les banques ont adhéré Baie, 23 (ATS)-La convention …
   Neue Zürcher Zeitung, 1982-10-01 | Die Vereinbarung über die Sorgfaltspflicht Inkrafttreten per 1. Oktober (ap) Am Freitag tritt die neue …
   Gazette de Lausanne, 1990-07-26 | BANQUES Blanchissage d'argent : les banques veulent maintenir le formulaire B La Commission fédérale des …
   Der Bund, 1987-09-28 | Banken mit aufgefrischten und griffigen Standesregeln Am 1. Oktober tritt die neue Sorgfaltspflichtvereinbarung bei der …

[12] voluntary credit limitation agreements  (n=192)
   Neue Zürcher Zeitung, 1980-03-11 | Weitgehende Aufhebung des Verzinsungsverbots Liberalisierung des Kapitalverkehrs zugunsten des Frankens Der Bundesrat und die Schweizerische …
   Gazette de Lausanne, 1981-01-10 | EN 1979,390 816 CRÉDITS À LA CONSOMMATION ONT ÉTÉ ACCORDÉS Mise sous tutelle du petit …
   Neue Zürcher Zeitung, 1996-01-31 | Der Konsumkreditmarkt im Widerstreit Mündige Bürger am Gängelband des Staates? Eine Vorort-Publikation zeigt, dass es …
   Journal de Genève, 1979-04-28 | CRÉDIT A LA CONSOMMATION Les réserves de l'UBS sur le projet de loi les débiteurs …
   Gazette de Lausanne, 1970-01-22 | Accord BNS banquiers Les taux d'accroissement des crédits abaissés La Banque Nationale Suisse et l'Association …

[13] registered shares and takeover defense  (n=67)
   Neue Zürcher Zeitung, 1988-04-30 | Sind vinkulierte Namenaktien börsenfähig? Das Bild gehört beinahe schon zur helvetischen Idylle: Just in dem …
   Neue Zürcher Zeitung, 1991-11-07 | Wenig transparente Schweizer Vinkulierungspraxis Zaghafte Ansätze einer Liberalisierung mrf. Die Vinkulierung von Namenaktien hat in …
   Neue Zürcher Zeitung, 1992-04-07 | Liberalisierter Erwerb von Schweizer Namenaktien Aufgehobene Empfehlung von 1961 ti. Der Verwaltungsrat der Schweizerischen Bankiervereinigung …
   Der Bund, 1987-04-16 | Kritik an Eintragungspraxis Aktienhandel erschwert sda. Die Vereinigung Schweizerischer Effektenbörsen setzt sich wie die Schweizerische …
   Gazette de Lausanne, 1991-04-12 | FINANCE Les gestionnaires de fortune en quête d'un statut Après cinq ans d'existence, l'Association suisse …

[14] european integration policy  (n=96)
   Journal de Genève, 1993-02-04 | Place financière genevoise : une carte à jouer race à Zurich Débat sur la placefinancière …
   Journal de Genève, 1992-10-19 | Les radicaux unis derrière l'EEE A plus de huit contre un, les délégués du Parti …
   Gazette de Lausanne, 1991-11-02 | EEE Réglementation bancaire Les banques suisses mises sur la touche Les membres de l'AELE sont …
   Journal de Genève, 1990-10-06 | -r Avertissement ferme deM. Bangemann aux pays de l'AELE Devant les banquiers réunis à Genève, …
   Gazette de Lausanne, 1989-06-21 | CEE Libéralisation bancaire et norme sur les initiés : satisfaction suisse La Suisse apprécie les …

[15] banking law revision  (n=220)
   Gazette de Lausanne, 1984-10-06 | JOURNÉE DES BANQUIERS L'ASB pour une révision rapide de la loi sur les banques Les …
   Journal de Genève, 1989-03-15 | Extension de la définition de la banque : les banquiers suisses favorables Baie (ATS).-Afin que …
   Neue Zürcher Zeitung, 1979-02-13 | Weitere Stellungnahmen Zentralkomitee des SIA (pd) Das Zentralkomitee des Schweizerischen Ingenieur-und Architekten-Vereins (SIA) empfiehlt Ablehnung …
   Journal de Genève, 1968-09-17 | ECONOMIE ET FINANCES Pour remédier à quelques faiblesses Faut-il retoucher la loi sur les banques …
   Neue Zürcher Zeitung, 1984-11-17 | Stimmen zum Konsumentenschutz CVP verlangt bessere Information (ap) Eine bessere Information der Konsumenten und einen …

[16] federal finance reform debates  (n=643)
   Journal de Genève, 1968-09-14 | Pourquoi adhérer au FMI ? Chaque année, dans la deuxième quinzaine de septembre, l'Association suisse …
   Gazette de Lausanne, 1985-02-14 | Place financière compétitive : propositions des banques BANQUES Zurich, 13 (ATS).-La place financière suisse ne …
   Journal de Genève, 1969-09-20 | Situation financière de la Confédération : ne pas « dramatiser » L'Association suisse des banquiers …
   Neue Zürcher Zeitung, 1966-05-17 | Der Ausbau des Instrumentariums der Nationalbank Veröffentlichung des Vorentwurfs + Bern, 17. Mai In jüngster …
   Neue Zürcher Zeitung, 1970-10-02 | Die Banken zum Münzgesetz Die Schweizerische Bankiervereinigung teilt mit: Gemäß Botschaft des Bundesrates zur Revision …

[17] mortgage interest rate policy  (n=127)
   Neue Zürcher Zeitung, 1975-02-22 | Weiterhin stabilisierte Hypothekarzinsen Ablehnung der Gesuche um Erhöhung auf 1. Juli 1975 Zy. Bern, 21. …
   Journal de Genève, 1994-01-11 | Les banques seront plus strictes dans l'octroi de crédits immobiliers La mise à disposition de …
   Gazette de Lausanne, 1963-10-02 | CONJONCTURE ET AMORTISSEMENT DE RÉCENTES statistiques des villes de Berne et Zurich nous apprennent que …
   Neue Zürcher Zeitung, 1980-08-27 | Gedämpftes Echo auf die Hypozinserhöhung Bern/Zürich, 26. Aug. (ddp) Die von der Schweizerischen Nationalbank (SNB) …
   Neue Zürcher Zeitung, 1980-08-26 | Erste Reaktionen «Hiobsbotschaft» für Hauseigentümer (ddp) Die Empfehlungen der Schweizerischen Bankiervereinigung für die Gestaltung der …

[18] bankers' day annual meeting  (n=260)
   Der Bund, 1973-10-07 | Für eine glaubwürdige Antlinflatlonspolitik Generalversammlung der Schweizerischen Bankiervereinigung in Basel a. k. Am Freitag fand …
   Gazette de Lausanne, 1976-03-05 | Conférence Eurofinas 1976 à Montreux • A Montreux se déroulera du 14 au 17 juin …
   Journal de Genève, 1957-09-30 | A BAiQ RAGAZ La Journée des banquiers (De notre envoyé spécial) (Dx.) C'est à Bad …
   Journal de Genève, 1963-10-07 | La journée des banquiers suisses $ Préoccupation majeure : l'effritement dufranc (De notre envoyé spécial.) …
   Journal de Genève, 1971-10-28 | -- •---•-••• •JE PENSE DONC JE VOTE Elections au Conseil national un spécialiste libéral pour …

[19] banks as employers  (n=128)
   Neue Zürcher Zeitung, 1997-09-06 | Kaum Streicheleinheiten für die Schweizer Banken Nüchterne Lageanalyse am Bankiertag in Bern Am schweizerischen Bankiertag …
   Neue Zürcher Zeitung, 1984-05-05 | Mehr sorgen für den Kanton 150 Jahre Berner Kantonalbank V Nach Ansicht von Bundesrat Stich …
   Der Bund, 1983-10-17 | Unsere Banken-ein wichtiger Teil im Wirtschafts-Puzzle. Unternehmen brauchen die Ban-\ Ken für Kredite. Zum Beispiel …
   Der Bund, 1982-11-26 | ; 1 I.;.. Reklame ^ /'^ ^ ¦ ~ | "lg In den letzten 20 …
   Journal de Genève, 1996-07-04 | Il y a encore 10 000 à 15 000 emplois bancaires de trop en Suisse …

Stability: how much does the partition change under other HDBSCAN settings? The adjusted Rand index (ARI) compares each alternative with the current clustering: 1 means identical, values near 0 mean unrelated.

In [25]:
rows = []
for mcs in [MIN_CLUSTER_SIZE // 2, MIN_CLUSTER_SIZE, MIN_CLUSTER_SIZE * 2]:
    for ms in [5, MIN_SAMPLES, MIN_SAMPLES * 2]:
        lab = HDBSCAN(min_cluster_size=mcs, min_samples=ms, cluster_selection_method=SELECTION, copy=True).fit_predict(Z)
        rows.append({"min_cluster_size": mcs, "min_samples": ms, "n_clusters": len(set(lab) - {-1}),
                     "unclustered_%": round((lab == -1).mean() * 100, 1),
                     "ARI_vs_current": round(adjusted_rand_score(dfc["cluster"], lab), 2)})
pd.DataFrame(rows)
Out[25]:
min_cluster_size min_samples n_clusters unclustered_% ARI_vs_current
0 30 5 60 24.6 0.25
1 30 50 28 44.3 0.79
2 30 100 11 40.0 0.36
3 61 5 30 27.9 0.39
4 61 50 20 41.6 1.00
5 61 100 10 39.9 0.36
6 122 5 10 13.5 0.20
7 122 50 8 28.3 0.32
8 122 100 8 42.3 0.33

10. Export an interactive page¶

The figure above is drawn by Plotly's own viewer. For the website we export a standalone HTML page instead, with a friendlier interface:

  • a hover box in the style of the site, coloured like the dot under the cursor: cluster name, newspaper, publication date and the first 15 words of the article;
  • a search bar that looks for words in those first 15 words and highlights the matching articles (case and accents are ignored, so zurich finds Zürich; several words must all be present);
  • two sliders for the opacity and the size of the dots, to read dense or sparse regions.

The page is one self-contained file: the data, the code and the styles are all inside it, and it needs no server. Only the first 15 words of each article are written into it. The cell stops with an error if any snippet is longer, and output/ never receives the rest of the text.

Set PLOTLY_JS = "inline" to embed the Plotly library too (about 4.5 MB more, but the page then works without an internet connection).

In [26]:
PLOTLY_JS = "cdn"            # "cdn": the library is loaded from plotly.com | "inline": embedded in the file
START_OPACITY, START_SIZE = 0.7, 4

PAGE = r"""<!doctype html>
<!-- (c) 2026 Arthur Michelet, CC BY-NC 4.0 (https://creativecommons.org/licenses/by-nc/4.0/).
     The article openings (first 15 words) shown on hover are quotations; they remain the property of their publishers. -->
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>__TITLE__</title>
<link rel="preconnect" href="https://fonts.googleapis.com"><link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
<link href="https://fonts.googleapis.com/css2?family=IBM+Plex+Mono:wght@500;700&family=Google+Sans+Flex&display=swap" rel="stylesheet">
__PLOTLY__
<style>
  :root { --ink: #333333; --paper: #ffffff; --mint: #C0F3E2; --green: #1D9E75; --deep: #085041; --orange: #FF4D14; }
  * { box-sizing: border-box; }
  html, body { margin: 0; height: 100%; background: var(--paper); color: var(--ink);
               font: 14px/1.4 'Google Sans Flex', 'Helvetica Neue', Arial, sans-serif; }
  .wrap { display: flex; flex-direction: column; gap: 10px; padding: 12px;
          width: 100% !important; height: 100vh !important; min-height: 460px; }
  .bar { display: flex; flex-wrap: wrap; align-items: center; gap: 10px 22px; }
  .bar label, .count { font: 500 12px 'IBM Plex Mono', Menlo, monospace; }
  .field { display: flex; align-items: center; gap: 8px; }
  .search { flex: 1 1 260px; max-width: 460px; }
  .search input { flex: 1; min-width: 0; padding: 6px 10px; font: inherit; color: var(--ink);
                  background: var(--paper); border: 2px solid var(--ink); border-radius: 2px; }
  .search input:focus-visible, button:focus-visible, input[type=range]:focus-visible { outline: 3px solid var(--green); outline-offset: 2px; }
  button { padding: 5px 10px; font: 500 12px 'IBM Plex Mono', Menlo, monospace; color: var(--ink);
           background: var(--paper); border: 2px solid var(--ink); border-radius: 2px; cursor: pointer; }
  button:hover { background: var(--mint); }
  input[type=range] { width: 110px; accent-color: var(--green); }
  output { min-width: 2.6em; font: 500 12px 'IBM Plex Mono', Menlo, monospace; }
  .count { margin-left: auto; color: rgba(51, 51, 51, .75); }
  .plotbox { position: relative; flex: 1 1 auto; min-height: 0; border: 2px solid var(--ink); border-radius: 2px; }
  .plot { position: absolute; inset: 0; }
  .caption { font-size: 11px; color: rgba(51, 51, 51, .75); }

  /* the hover box: coloured like the dot, with the opening of the article on a white inset */
  #tip { position: fixed; z-index: 10; left: 0; top: 0; display: none; width: max-content; max-width: 330px;
         padding: 8px; pointer-events: none; border: 2px solid var(--ink); border-radius: 2px;
         box-shadow: -4px 4px 0 var(--ink); }
  #tip .cluster { margin: 0 0 6px; font: 700 13px/1.25 'IBM Plex Mono', Menlo, monospace; }
  #tip .meta { margin: 0 0 6px; font-size: 12.5px; }
  #tip .meta b { font-weight: 700; }
  #tip .snippet { margin: 0; padding: 6px 8px; font-size: 13px; font-style: italic; color: var(--ink);
                  background: var(--paper); border: 2px solid var(--ink); border-radius: 2px; }
</style>
</head>
<body>
<div class="wrap" style="height:__PAGE_H__px; width:__PAGE_W__px;">
  <div class="bar">
    <div class="field search">
      <label for="q">search</label>
      <input id="q" type="search" placeholder="words in the first __N__ words" autocomplete="off" spellcheck="false">
      <button id="clear" type="button" title="clear the search">clear</button>
    </div>
    <div class="field"><label for="op">opacity</label><input id="op" type="range" min="0.05" max="1" step="0.05" value="__OP__"><output id="op-out"></output></div>
    <div class="field"><label for="sz">size</label><input id="sz" type="range" min="1" max="14" step="0.5" value="__SZ__"><output id="sz-out"></output></div>
    <div class="count" id="count" aria-live="polite"></div>
  </div>
  <div class="plotbox"><div id="umap" class="plot"></div></div>
  <div class="caption">__CAPTION__</div>
</div>
<div id="tip" role="tooltip"><p class="cluster"></p><p class="meta"></p><p class="snippet"></p></div>

<script>
(function () {
  const FIG = __FIG__;
  const META = __META__;
  const YEAR_STYLE = __YEAR_STYLE__;
  const Y0 = __Y0__, Y1 = __Y1__, DIM = 0.05;
  const MONTHS = ['January','February','March','April','May','June','July','August','September','October','November','December'];

  const gd = document.getElementById('umap'), tip = document.getElementById('tip');
  const q = document.getElementById('q'), op = document.getElementById('op'), sz = document.getElementById('sz');
  const count = document.getElementById('count');
  const N = META.yr.length;
  const fmt = new Intl.NumberFormat('en');
  const plain = s => s.normalize('NFD').replace(/[\u0300-\u036f]/g, '').toLowerCase();
  const haystack = META.sn.map(plain);                      // searchable text: lower case, accents removed
  let matching = null;                                      // null: no search, else the matching row numbers

  function niceDate(iso) {
    const p = iso.split('-');
    return p.length === 3 ? (+p[2]) + ' ' + MONTHS[+p[1] - 1] + ' ' + p[0] : 'date unknown';
  }

  function showTip(i, ev) {
    const style = YEAR_STYLE[Math.min(Math.max(META.yr[i], Y0), Y1)];
    tip.style.background = style[0]; tip.style.color = style[1];
    tip.querySelector('.cluster').textContent = META.clusters[META.ci[i]] || '';
    tip.querySelector('.cluster').style.display = META.clusters[META.ci[i]] ? '' : 'none';
    const meta = tip.querySelector('.meta'); meta.textContent = '';
    const b = document.createElement('b'); b.textContent = META.papers[META.pi[i]];
    meta.append(b, ' · ' + niceDate(META.dt[i]));
    tip.querySelector('.snippet').textContent = META.sn[i];
    tip.style.display = 'block';
    const w = tip.offsetWidth, h = tip.offsetHeight, m = 14;
    let x = ev.clientX + m, y = ev.clientY + m;
    if (x + w > innerWidth - 4) x = ev.clientX - w - m;
    if (y + h > innerHeight - 4) y = ev.clientY - h - m;
    tip.style.left = Math.max(4, x) + 'px'; tip.style.top = Math.max(4, y) + 'px';
  }
  const hideTip = () => { tip.style.display = 'none'; };

  function styleDots() {                                    // opacity and size sliders
    const o = +op.value, s = +sz.value;
    document.getElementById('op-out').textContent = o.toFixed(2);
    document.getElementById('sz-out').textContent = s;
    Plotly.restyle(gd, {'marker.opacity': [matching ? DIM : o], 'marker.size': [s]}, [0]);
    Plotly.restyle(gd, {'marker.opacity': [o < 0.6 ? 1 : o], 'marker.size': [s + 1]}, [1]);
  }

  function search() {                                       // every term must appear in the first words
    const terms = plain(q.value).split(/\s+/).filter(Boolean);
    if (!terms.length) {
      matching = null;
      Plotly.restyle(gd, {x: [[]], y: [[]], 'marker.color': [[]], customdata: [[]]}, [1]);
      count.textContent = fmt.format(N) + ' articles';
    } else {
      matching = [];
      for (let i = 0; i < N; i++) if (terms.every(t => haystack[i].includes(t))) matching.push(i);
      Plotly.restyle(gd, {
        x: [matching.map(i => X[i])], y: [matching.map(i => Y[i])],
        'marker.color': [matching.map(i => META.yr[i])], customdata: [matching],
      }, [1]);
      count.textContent = fmt.format(matching.length) + ' of ' + fmt.format(N) + ' articles match';
    }
    Plotly.restyle(gd, {hoverinfo: [matching ? 'skip' : 'none']}, [0]);   // while searching, only matches answer the pointer
    hideTip(); styleDots();
  }

  // plotly.py stores number arrays as base64 ({dtype, bdata}); read the coordinates back as plain numbers
  const DTYPES = {f8: Float64Array, f4: Float32Array, i4: Int32Array, i2: Int16Array, i1: Int8Array, u4: Uint32Array, u2: Uint16Array, u1: Uint8Array};
  function numbers(a) {
    if (Array.isArray(a)) return a;
    const raw = atob(a.bdata), bytes = new Uint8Array(raw.length);
    for (let k = 0; k < raw.length; k++) bytes[k] = raw.charCodeAt(k);
    return Array.from(new DTYPES[a.dtype](bytes.buffer));
  }
  const X = numbers(FIG.data[0].x), Y = numbers(FIG.data[0].y);
  let timer;
  const config = {responsive: true, displaylogo: false, scrollZoom: true,
                  modeBarButtonsToRemove: ['toImage', 'sendDataToCloud', 'sendChartToCloud', 'select2d', 'lasso2d', 'autoScale2d']};
  Plotly.newPlot(gd, FIG.data, FIG.layout, config).then(function () {
    count.textContent = fmt.format(N) + ' articles'; styleDots();
    gd.on('plotly_hover', function (ev) {
      const p = ev.points && ev.points[0];
      if (!p) return hideTip();
      showTip(p.curveNumber === 1 ? p.customdata : p.pointNumber, ev.event);
    });
    gd.on('plotly_unhover', hideTip);
    gd.addEventListener('mouseleave', hideTip);
  });

  op.addEventListener('input', styleDots);
  sz.addEventListener('input', styleDots);
  q.addEventListener('input', function () { clearTimeout(timer); timer = setTimeout(search, 180); });
  q.addEventListener('keydown', e => { if (e.key === 'Escape') { q.value = ''; search(); } });
  document.getElementById('clear').addEventListener('click', function () { q.value = ''; search(); q.focus(); });
})();
</script>
</body>
</html>
"""


def _js(obj):
    """JSON that is safe to place inside a <script> tag."""
    return (obj if isinstance(obj, str) else json.dumps(obj, ensure_ascii=False)) \
        .replace("<", "\\u003c").replace("\u2028", "\\u2028").replace("\u2029", "\\u2029")


def export_interactive_html(pts, anchors, path, plotly_js=PLOTLY_JS, opacity=START_OPACITY, size=START_SIZE,
                            title="UMAP of SBA coverage, 1940–1999", page_size=(1100, 810)):
    """Write the standalone interactive page. Only the first SNIPPET_WORDS words of each article go into it."""
    longest = max(len(s.removesuffix(" …").split()) for s in pts["snippet"])
    assert longest <= SNIPPET_WORDS, f"A snippet has {longest} words: at most {SNIPPET_WORDS} may be published"

    fig = make_figure(pts, anchors, size=size, opacity=opacity, interactive_export=True)
    papers, clusters = sorted(pts["paper"].unique()), sorted(pts["cluster_name"].unique())
    meta = {                                   # everything the hover box and the search need, per article
        "papers": papers, "clusters": clusters,
        "pi": pts["paper"].map({p: i for i, p in enumerate(papers)}).tolist(),
        "ci": pts["cluster_name"].map({c: i for i, c in enumerate(clusters)}).tolist(),
        "yr": pts["year"].tolist(), "dt": pts["date"].tolist(), "sn": pts["snippet"].tolist(),
    }
    library = (f"<script>{get_plotlyjs()}</script>" if plotly_js == "inline"
               else f'<script src="https://cdn.plot.ly/plotly-{get_plotlyjs_version()}.min.js" charset="utf-8"></script>')
    caption = f"{corpus_caption(len(pts))} Each article is shown by its first {SNIPPET_WORDS} words only."

    page = (PAGE
            .replace("__TITLE__", html.escape(title)).replace("__PLOTLY__", library)
            .replace("__PAGE_W__", str(page_size[0])).replace("__PAGE_H__", str(page_size[1]))
            .replace("__N__", str(SNIPPET_WORDS)).replace("__OP__", str(opacity)).replace("__SZ__", str(size))
            .replace("__CAPTION__", caption)
            .replace("__Y0__", str(Y0)).replace("__Y1__", str(Y1))
            .replace("__YEAR_STYLE__", _js(YEAR_STYLE))
            .replace("__META__", _js(meta)).replace("__FIG__", _js(to_json(fig))))
    Path(path).write_text(page, encoding="utf-8")
    print(f"{path}: {Path(path).stat().st_size / 1e6:.1f} MB, {len(pts):,} articles")


export_interactive_html(pts, anchors, HTML_PATH)
output/umap-labelled-map.html: 1.1 MB, 6,153 articles

Open output/umap-labelled-map.html in a browser to try it. Hover over a dot, type a word in the search bar, and move the sliders.

To publish the map, copy the HTML file to your website and add your own link to this notebook. Before pushing anything to a public repository, check that only output/ and the notebook are included: data/ holds the full article texts and must stay private (it is excluded from Git by the .gitignore created in stage 1).

Licence¶

Text, figures and narration: CC BY-NC 4.0, code: MIT (c) 2026 Arthur Michelet. The article openings shown in the map are short quotations from newspapers retrieved through Impresso; they remain the property of their publishers and are not covered by this licence.