Everything Majors Are Worried About, in a Scatterplot
Every year, a few hundred majors and lieutenant-commanders graduate from the Joint Command and Staff Programme (JCSP) at the Canadian Forces College (CFC) in Toronto. Each of them writes at least one paper, and the College posts the papers online, where they sit in a database (…that seems to be broken, as of writing).
JCSP is a formative year. The graduates go straight into jobs where they influence policy, and a subset will write it as they climb into the executive ranks. Some of the bold ideas in these papers do come to fruition, so the entire repo is a bit of a crystal ball. I feel like Nostradamus spelunking through them.
Anyways, I had an idea: trace force development efforts by reading what has been written en masse. I am working on my own innovation efforts within my command, but what have these tenured dinosaurs put forward that might crest the horizon soon? Perhaps the answer was sitting in that busted database.
So I downloaded all the papers. Of 1,248 listings, I recovered 1,081 full texts, 8.8 million words.
This is what five years of Canadian staff college papers look like from above.
Every dot is a paper, placed by what it argues about, so the Arctic papers huddle together and the shipbuilding papers huddle somewhere else. The rest of this post is about how that picture was made.
Last time, I proposed reading performance reports with a language model. This time the tools are pretty humble: counting words, a small embedding model, and a clustering algorithm. Simple, but mighty tools!
Acronyms
- C2
- Command and control
- CAF
- Canadian Armed Forces
- CDS
- Chief of the Defence Staff
- CFC
- Canadian Forces College
- CJFC
- Canadian Joint Force Command (2025)
- CJOC
- Canadian Joint Operations Command
- DM
- Deputy Minister
- FD
- Force development
- HDBSCAN
- Hierarchical density-based clustering
- JCSP
- Joint Command and Staff Programme
- LLM
- Large language model
- MDO
- Multi-domain operations
- MDS
- Master of Defence Studies
- NORAD
- North American Aerospace Defense Command
- NSP
- National Security Programme
- ONSAF
- Our North, Strong and Free (2024 defence policy)
- SOF
- Special operations forces
- SSE
- Strong, Secure, Engaged (2017 defence policy)
- UMAP
- Uniform manifold approximation and projection
The corpus
JCSP students produce three kinds of papers:
- Service Paper. Short, about 3,600 words, addressed to a named commander, written in the first term. It answers what should we buy, or reorganize?
- Solo Flight. A free-choice essay of about 6,100 words, written later in the year. It answers what’s wrong, and where are we?
- MDS thesis. Optional, around 25,000 words, for the students who take the degree. It touches on just about everything.
The College’s index lists papers by year, with author, title, and a link to the PDF 1. Or it would, if it worked.
So I checked whether archive.org had, y’know, archived it.
Success! Nothing you post is ever truly gone. The snapshot’s filters were broken, though, so I paged through all 507 pages of it with a throwaway Python script: httpx2 to fetch, bs4 to parse the listings, and pypdf to pull text out of the PDFs.
The archive had a rough go of the 2025 cohort. Of 1,248 listings, 170 came back as dead links, and 160 of those are 2025. The most recent year is fifteen papers throughout, which is why it is starred and dashed in every chart. I’ll refresh the dataset when the live index is working again.
| 2021 | 2022 | 2023 | 2024 | 2025* | All | |
|---|---|---|---|---|---|---|
| Solo Flight | 82 | 177 | 72 | 147 | 5 | 483 |
| Service Paper | 113 | 124 | 96 | 97 | 8 | 438 |
| MDS | 92 | 18 | 19 | 27 | 2 | 158 |
| NSP | – | – | – | 2 | – | 2 |
| Texts | 287 | 319 | 187 | 273 | 15 | 1,081 |
Methodology
Three passes over the text: count, embed, cluster. Each is a few dozen lines of Python.
Count. For each paper, it was trivial to make a word cloud and pull out its common themes, which gave me a baseline set of topics.
from collections import Counter
Counter(text.lower().split()).most_common(20) A paper that says “Arctic” once in a footnote is likely not about the Arctic, so the headline figure in most charts is the share of papers with three or more hits for rigour.
Embedding. This is a bit of an abstract concept, for non-technical folks . Counting only finds the words I looked for, but an embedding model finds the rest. It turns a passage into a point in space. Passages about the same thing land “near” each other, whatever words or language they use 2. I embedded each paper in 200-word chunks and averaged them, so every paper is one point.
model = SentenceTransformer("paraphrase-multilingual-MiniLM-L12-v2")
chunks = [" ".join(words[i : i + 200]) for i in range(0, min(len(words), 6_000), 200)]
vectors[paper_id] = model.encode(chunks).mean(axis=0) # 384 numbers per paper Down-project and cluster. Those points live in 384 dimensions, too many to cluster and impossible to visualize. UMAP squashes them down to five while keeping neighbours as neighbours, and HDBSCAN finds the dense regions and leaves strays as outliers 3.
Five dimensions were chosen because HDBSCAN works by density, and in high dimensions every point is about the same distance from every other, so there is nothing dense to find. Two dimensions, which the map at the top uses, crushes unrelated papers together just to fit on a page and makes a pretty visualization.
low = umap.UMAP(n_components=5, metric="cosine").fit_transform(vectors)
cluster = hdbscan.HDBSCAN(min_cluster_size=15).fit_predict(low) Nineteen clusters fell out of this process. In sequence, each cluster was named by reading each one’s characteristic words and neighboring papers.
The reference war changed
Data
| 2021 | 2022 | 2023 | 2024 | 2025 | |
|---|---|---|---|---|---|
| Ukraine | 17.4 | 33.2 | 38 | 46.2 | 40 |
| Afghanistan | 33.8 | 26.3 | 28.3 | 24.2 | 26.7 |
In 2021, a third of papers cited Afghanistan and a sixth cited Ukraine. By 2022 the lines had crossed. By 2024 Ukraine appears in nearly half the corpus.
It also changed how capability gets argued. Density of uncrewed-systems vocabulary nearly triples in 2023, from about five hits per ten thousand words to thirteen, and stays there.
A word’s life cycle
Data
| 2021 | 2022 | 2023 | 2024 | 2025 | |
|---|---|---|---|---|---|
| Retention | 29.3 | 30.4 | 29.9 | 39.6 | 46.7 |
| Reconstitution | 2.1 | 12.5 | 23.5 | 14.7 | 20 |
| Talent management | 2.4 | 2.8 | 2.1 | 6.6 | 0 |
“Reconstitution” is the cleanest case study in the corpus. It appears in 2 percent of papers in 2021, 12 percent in 2022, and peaks at 23 percent in 2023, the first full writing year after the CDS and DM directive of October 2022 4. Then, it fell back to 15 percent in 2024 while the things it named keep rising. Retention goes from 29 to 40 percent of papers. “Talent management,” essentially absent before, surfaces more in 2024.
Of course, we are now seeing “talent management boards”, internally! Crystal ball, indeed.
The programme picks the subject
| Theme | Solo Flight n = 483 | Service Paper n = 438 | MDS n = 158 |
|---|---|---|---|
| Personnel & reconstitution | 59 | 38 | 89 |
| Strategic context | 51 | 28 | 58 |
| Modernization & procurement | 32 | 45 | 62 |
| Emerging tech | 31 | 38 | 56 |
| Continental defence & Arctic | 24 | 16 | 38 |
| Readiness & force generation | 23 | 22 | 49 |
| Pan-domain / joint C2 | 22 | 36 | 44 |
| Corps restructuring / Force 2025 | 9 | 13 | 17 |
Modernization and joint C2 appear substantively in 36 to 45 percent of Service Papers but only 22 to 32 percent of Solo Flight essays. Personnel appears in 59 percent of Solo Flight essays, strategic context in 51 percent. Theses hit most themes because they’re long enough to.
Data
| 2021 | 2022 | 2023 | 2024 | 2025 | |
|---|---|---|---|---|---|
| COVID / pandemic | 31.4 | 36.1 | 25.7 | 23.1 | 13.3 |
| Culture change | 13.9 | 17.2 | 15.5 | 9.9 | 6.7 |
| Arbour report | 0.3 | 1.9 | 8 | 5.1 | 6.7 |
One more from the counts. Culture and conduct vocabulary peaks in 2022, after the Arbour report 5, and halves by 2024. The pandemic follows the same curve: 36 percent of papers in 2022, 23 percent in 2024.
What the topic model adds
The keyword counts only find the forty things I thought to look for, but the embeddings find everything else. The map at the top of this post is those embeddings squashed to two dimensions, and the clusters below are what HDBSCAN found in five.
- Army structure & land modernization cavalry · mdo · army reserve · leopard · combat team131
- Air power & uncrewed aviation uas · uav · airpower · tactical aviation · rpas124
- Personnel: HR, retention, families family support · pension · relocation · millennials · retention strategy105
- Arctic & continental defence arctic security · northwest passage · lackenbauer · arctic council · norad modernization92
- Navy: shipbuilding & maritime capability shipbuilding · csc · victoria class · leadmark · asw67
- Leadership & culture change hateful · radicalization · emotional intelligence · toxic · mentoring59
- Expeditionary & peace opsFrench mali · peacekeepers · peace operations · minusma · djibouti52
- Defence enterprise: procurement, domestic ops, force size domestic operations · project approval · emergency management · kpmg51
- Gender, diversity & inclusion employment equity · women peace · gender-based · masculinity47
- Grey-zone, hybrid & SOF strategic culture · unrestricted warfare · grey zone · way of war47
- National security & grand strategy securitization · security culture · grand strategy · migrants45
- Cyber & intelligence stuxnet · malware · cse · blockchain44
- Institutional / HRFrench connaissances · organisation · information44
- Digital transformation & logistics digital transformation · outsourcing · digital literacy · additive manufacturing41
- Russia / NATO Europe russie · baltic · putin · finland34
- China & Indo-Pacific brics · prc · belt and road · south china sea31
- AI & autonomous weapons autonomous weapons · robots · lethal autonomous · killer28
- Space domain space domain · space operations · outer space · space policy22
- Health, fitness & resilience fitness · résilience · santé · nutrition17
A few things the counts add to that picture:
- Steady backdrops. Cyber, space, climate, and China barely move. China and the Indo-Pacific sit in roughly four papers in ten every year.
- New arrivals in 2024. Gaza and the Red Sea, Taiwan, and generative AI all jump from near zero; I imagine trends in 2025 and 2026 will escalate this.
- Readiness, maybe. Three-quarters of the fifteen 2025 papers mention it, against a third to two-fifths in every earlier year. Real, or a sampling error based on which 2025 PDFs survived? The 160 missing papers would settle it… and I don’t have them.
Regardless, the College publishes these papers so they’ll be read— this is just one way to read all of them!
Update - 8 Sept 26: Letting a local LLM name the topics!
Toponymy 6 is a wholly 🇨🇦 effort— from the same team behind UMAP/HDBSCAN— and is an LLM-powered topic-labeling tool. It clusters the same two-dimensional map at several densities, so each broad region contains finer ones, then hands every cluster’s keyphrases and a few representative papers to an LLM before querying for a representative name.
clusterer = ToponymyClusterer(min_clusters=4, base_min_cluster_size=12)
llm = LiteLLMNamer(
model="openai/google/gemma-4-12b-qat",
api_base="http://localhost:1234/v1",
)
model = Toponymy(llm, embedder, clusterer, object_description="staff college papers")
model.fit(texts, embedding_vectors=vectors, clusterable_vectors=map_xy) Thank you, Dr. T.J., for the Toponomy suggestion!
- The Canadian Forces College papers index,
cfc.forces.gc.ca, via Wayback Machine captures from September 2025. Full text extracted from the PDFs.↩ paraphrase-multilingual-MiniLM-L12-v2, a 118M-parameter sentence-transformer trained on parallel data in 50 languages, so English and French papers land in the same space.↩- Both steps run inside BERTopic, which also supplies each cluster’s characteristic words via c-TF-IDF. HDBSCAN left 132 papers as outliers; for the prevalence figures they were assigned to the nearest cluster by embedding.↩
- CDS/DM Directive for CAF Reconstitution, 6 October 2022.↩
- Louise Arbour, Report of the Independent External Comprehensive Review of the Department of National Defence and the Canadian Armed Forces, May 2022.↩
- Tutte Institute for Mathematics and Computing, Toponymy 0.5.4. Layers are clustered on a two-dimensional layout rather than the five-dimensional one behind the nineteen clusters, which is coarser (but keeps every topic contiguous on screen!). Names were generated by Gemma 4 12B-QAT running locally in LM Studio and are entirely unedited.↩