grimblog

Everything Majors Are Worried About, in a Scatterplot

Every year, a few hundred majors and lieutenant-commanders graduate from the Joint Command and Staff Programme (JCSP) at the Canadian Forces College (CFC) in Toronto. Each of them writes at least one paper, and the College posts the papers online, where they sit in a database (…that seems to be broken, as of writing).

JCSP is a formative year. The graduates go straight into jobs where they influence policy, and a subset will write it as they climb into the executive ranks. Some of the bold ideas in these papers do come to fruition, so the entire repo is a bit of a crystal ball. I feel like Nostradamus spelunking through them.

Anyways, I had an idea: trace force development efforts by reading what has been written en masse. I am working on my own innovation efforts within my command, but what have these tenured dinosaurs put forward that might crest the horizon soon? Perhaps the answer was sitting in that busted database.

So I downloaded all the papers. Of 1,248 listings, I recovered 1,081 full texts, 8.8 million words.

This is what five years of Canadian staff college papers look like from above.

Colour
Year
1,081 of 1,081
Every dot is a paper placed by a 2-D UMAP of its text embedding; nearby papers argue about similar things. Axes are arbitrary. Scroll or pinch to zoom, drag to pan. Click a dot to pin its card, then click the title to open the PDF.

Every dot is a paper, placed by what it argues about, so the Arctic papers huddle together and the shipbuilding papers huddle somewhere else. The rest of this post is about how that picture was made.

Last time, I proposed reading performance reports with a language model. This time the tools are pretty humble: counting words, a small embedding model, and a clustering algorithm. Simple, but mighty tools!

Acronyms
C2
Command and control
CAF
Canadian Armed Forces
CDS
Chief of the Defence Staff
CFC
Canadian Forces College
CJFC
Canadian Joint Force Command (2025)
CJOC
Canadian Joint Operations Command
DM
Deputy Minister
FD
Force development
HDBSCAN
Hierarchical density-based clustering
JCSP
Joint Command and Staff Programme
LLM
Large language model
MDO
Multi-domain operations
MDS
Master of Defence Studies
NORAD
North American Aerospace Defense Command
NSP
National Security Programme
ONSAF
Our North, Strong and Free (2024 defence policy)
SOF
Special operations forces
SSE
Strong, Secure, Engaged (2017 defence policy)
UMAP
Uniform manifold approximation and projection

The corpus

JCSP students produce three kinds of papers:

  • Service Paper. Short, about 3,600 words, addressed to a named commander, written in the first term. It answers what should we buy, or reorganize?
  • Solo Flight. A free-choice essay of about 6,100 words, written later in the year. It answers what’s wrong, and where are we?
  • MDS thesis. Optional, around 25,000 words, for the students who take the degree. It touches on just about everything.

The College’s index lists papers by year, with author, title, and a link to the PDF 1. Or it would, if it worked.

So I checked whether archive.org had, y’know, archived it.

Success! Nothing you post is ever truly gone. The snapshot’s filters were broken, though, so I paged through all 507 pages of it with a throwaway Python script: httpx2 to fetch, bs4 to parse the listings, and pypdf to pull text out of the PDFs.

The archive had a rough go of the 2025 cohort. Of 1,248 listings, 170 came back as dead links, and 160 of those are 2025. The most recent year is fifteen papers throughout, which is why it is starred and dashed in every chart. I’ll refresh the dataset when the live index is working again.

20212022202320242025*All
Solo Flight82177721475483
Service Paper11312496978438
MDS921819272158
NSP22
Texts287 319 187 273 15 1,081
Recovered full texts by programme and year. *2025 is 15 papers.

Methodology

Three passes over the text: count, embed, cluster. Each is a few dozen lines of Python.

Count. For each paper, it was trivial to make a word cloud and pull out its common themes, which gave me a baseline set of topics.

from collections import Counter

Counter(text.lower().split()).most_common(20)

A paper that says “Arctic” once in a footnote is likely not about the Arctic, so the headline figure in most charts is the share of papers with three or more hits for rigour.

Embedding. This is a bit of an abstract concept, for non-technical folks . Counting only finds the words I looked for, but an embedding model finds the rest. It turns a passage into a point in space. Passages about the same thing land “near” each other, whatever words or language they use 2. I embedded each paper in 200-word chunks and averaged them, so every paper is one point.

model = SentenceTransformer("paraphrase-multilingual-MiniLM-L12-v2")

chunks = [" ".join(words[i : i + 200]) for i in range(0, min(len(words), 6_000), 200)]
vectors[paper_id] = model.encode(chunks).mean(axis=0)  # 384 numbers per paper

Down-project and cluster. Those points live in 384 dimensions, too many to cluster and impossible to visualize. UMAP squashes them down to five while keeping neighbours as neighbours, and HDBSCAN finds the dense regions and leaves strays as outliers 3.

Five dimensions were chosen because HDBSCAN works by density, and in high dimensions every point is about the same distance from every other, so there is nothing dense to find. Two dimensions, which the map at the top uses, crushes unrelated papers together just to fit on a page and makes a pretty visualization.

low = umap.UMAP(n_components=5, metric="cosine").fit_transform(vectors)
cluster = hdbscan.HDBSCAN(min_cluster_size=15).fit_predict(low)

Nineteen clusters fell out of this process. In sequence, each cluster was named by reading each one’s characteristic words and neighboring papers.

The reference war changed

0102030405020212022202320242025*UkraineAfghanistan
2024 Ukraine 46.2 Afghanistan 24.2
Data
20212022202320242025
Ukraine17.433.23846.240
Afghanistan33.826.328.324.226.7
Papers citing each conflict, any mention. % of papers. *2025 is 15 papers.

In 2021, a third of papers cited Afghanistan and a sixth cited Ukraine. By 2022 the lines had crossed. By 2024 Ukraine appears in nearly half the corpus.

It also changed how capability gets argued. Density of uncrewed-systems vocabulary nearly triples in 2023, from about five hits per ten thousand words to thirteen, and stays there.

A word’s life cycle

0102030405020212022202320242025*Reconstitution DirectiveRetentionReconstitutionTalent management
2024 Retention 39.6 Reconstitution 14.7 Talent management 6.6
Data
20212022202320242025
Retention29.330.429.939.646.7
Reconstitution2.112.523.514.720
Talent management2.42.82.16.60
Personnel vocabulary, any mention. % of papers. *2025 is 15 papers.

“Reconstitution” is the cleanest case study in the corpus. It appears in 2 percent of papers in 2021, 12 percent in 2022, and peaks at 23 percent in 2023, the first full writing year after the CDS and DM directive of October 2022 4. Then, it fell back to 15 percent in 2024 while the things it named keep rising. Retention goes from 29 to 40 percent of papers. “Talent management,” essentially absent before, surfaces more in 2024.

Of course, we are now seeing “talent management boards”, internally! Crystal ball, indeed.

The programme picks the subject

ThemeSolo Flight n = 483Service Paper n = 438MDS n = 158
Personnel & reconstitution 59 38 89
Strategic context 51 28 58
Modernization & procurement 32 45 62
Emerging tech 31 38 56
Continental defence & Arctic 24 16 38
Readiness & force generation 23 22 49
Pan-domain / joint C2 22 36 44
Corps restructuring / Force 2025 9 13 17
Share of each programme's papers with three or more mentions of any term in the theme.

Modernization and joint C2 appear substantively in 36 to 45 percent of Service Papers but only 22 to 32 percent of Solo Flight essays. Personnel appears in 59 percent of Solo Flight essays, strategic context in 51 percent. Theses hit most themes because they’re long enough to.

081624324020212022202320242025*COVID / pandemicCulture changeArbour report
2024 COVID / pandemic 23.1 Culture change 9.9 Arbour report 5.1
Data
20212022202320242025
COVID / pandemic31.436.125.723.113.3
Culture change13.917.215.59.96.7
Arbour report0.31.985.16.7
Topics that peaked, any mention. % of papers. *2025 is 15 papers.

One more from the counts. Culture and conduct vocabulary peaks in 2022, after the Arbour report 5, and halves by 2024. The pandemic follows the same curve: 36 percent of papers in 2022, 23 percent in 2024.

What the topic model adds

The keyword counts only find the forty things I thought to look for, but the embeddings find everything else. The map at the top of this post is those embeddings squashed to two dimensions, and the clusters below are what HDBSCAN found in five.

  1. Army structure & land modernization cavalry · mdo · army reserve · leopard · combat team
    131
  2. Air power & uncrewed aviation uas · uav · airpower · tactical aviation · rpas
    124
  3. Personnel: HR, retention, families family support · pension · relocation · millennials · retention strategy
    105
  4. Arctic & continental defence arctic security · northwest passage · lackenbauer · arctic council · norad modernization
    92
  5. Navy: shipbuilding & maritime capability shipbuilding · csc · victoria class · leadmark · asw
    67
  6. Leadership & culture change hateful · radicalization · emotional intelligence · toxic · mentoring
    59
  7. Expeditionary & peace opsFrench mali · peacekeepers · peace operations · minusma · djibouti
    52
  8. Defence enterprise: procurement, domestic ops, force size domestic operations · project approval · emergency management · kpmg
    51
  9. Gender, diversity & inclusion employment equity · women peace · gender-based · masculinity
    47
  10. Grey-zone, hybrid & SOF strategic culture · unrestricted warfare · grey zone · way of war
    47
  11. National security & grand strategy securitization · security culture · grand strategy · migrants
    45
  12. Cyber & intelligence stuxnet · malware · cse · blockchain
    44
  13. Institutional / HRFrench connaissances · organisation · information
    44
  14. Digital transformation & logistics digital transformation · outsourcing · digital literacy · additive manufacturing
    41
  15. Russia / NATO Europe russie · baltic · putin · finland
    34
  16. China & Indo-Pacific brics · prc · belt and road · south china sea
    31
  17. AI & autonomous weapons autonomous weapons · robots · lethal autonomous · killer
    28
  18. Space domain space domain · space operations · outer space · space policy
    22
  19. Health, fitness & resilience fitness · résilience · santé · nutrition
    17
Nineteen BERTopic clusters on 1,081 papers, ordered by size. Sparkline is each cluster's share of the year's papers, 2021–2024, on a common scale; 2025 is omitted at 15 papers.

A few things the counts add to that picture:

  • Steady backdrops. Cyber, space, climate, and China barely move. China and the Indo-Pacific sit in roughly four papers in ten every year.
  • New arrivals in 2024. Gaza and the Red Sea, Taiwan, and generative AI all jump from near zero; I imagine trends in 2025 and 2026 will escalate this.
  • Readiness, maybe. Three-quarters of the fifteen 2025 papers mention it, against a third to two-fifths in every earlier year. Real, or a sampling error based on which 2025 PDFs survived? The 160 missing papers would settle it… and I don’t have them.

Regardless, the College publishes these papers so they’ll be read— this is just one way to read all of them!

Update - 8 Sept 26: Letting a local LLM name the topics!

Toponymy 6 is a wholly 🇨🇦 effort— from the same team behind UMAP/HDBSCAN— and is an LLM-powered topic-labeling tool. It clusters the same two-dimensional map at several densities, so each broad region contains finer ones, then hands every cluster’s keyphrases and a few representative papers to an LLM before querying for a representative name.

clusterer = ToponymyClusterer(min_clusters=4, base_min_cluster_size=12)
llm = LiteLLMNamer(
    model="openai/google/gemma-4-12b-qat", 
    api_base="http://localhost:1234/v1",
)

model = Toponymy(llm, embedder, clusterer, object_description="staff college papers")
model.fit(texts, embedding_vectors=vectors, clusterable_vectors=map_xy)
Colour
Topics
Year
1,081 of 1,081

Thank you, Dr. T.J., for the Toponomy suggestion!


  1. The Canadian Forces College papers index, cfc.forces.gc.ca, via Wayback Machine captures from September 2025. Full text extracted from the PDFs.
  2. paraphrase-multilingual-MiniLM-L12-v2, a 118M-parameter sentence-transformer trained on parallel data in 50 languages, so English and French papers land in the same space.
  3. Both steps run inside BERTopic, which also supplies each cluster’s characteristic words via c-TF-IDF. HDBSCAN left 132 papers as outliers; for the prevalence figures they were assigned to the nearest cluster by embedding.
  4. CDS/DM Directive for CAF Reconstitution, 6 October 2022.
  5. Louise Arbour, Report of the Independent External Comprehensive Review of the Department of National Defence and the Canadian Armed Forces, May 2022.
  6. Tutte Institute for Mathematics and Computing, Toponymy 0.5.4. Layers are clustered on a two-dimensional layout rather than the five-dimensional one behind the nineteen clusters, which is coarser (but keeps every topic contiguous on screen!). Names were generated by Gemma 4 12B-QAT running locally in LM Studio and are entirely unedited.

← All posts