My Peace Study

2026-08-09

About the Data Used Here: How It Was Counted, and How It Is Skewed

How the counts are made, and the four skews in this data (subjective extraction / it is promotional copy / it passed censorship / the people who could publish were a minority).

Numbers show up all over this site. "Of 883 notes, 5 mention censorship." "War words 757 / reconstruction words 858." Numbers like these look objective.

But these numbers did not come out of a statistical survey. I read things, I picked things out, and afterward I counted what I had picked. What was counted, and how, and where it is skewed — I am setting all of that down here.

What I Read

I searched the National Diet Library Digital Collections for books published between June 1, 1945 and September 30, 1946, and took as my subject the ones whose text I could actually read, either published online or sent to libraries. I read every preface and afterword.

As I read, I copied out whatever struck me as "here is what he really thinks," "I just like this," or "this is about the war," and made notes of it. Those notes are what the numbers on this site are built from.

The chronology page alone is different: it is a record of "what somebody was doing on that day," gathered from diaries, memoirs, magazine articles, and all sorts of other places.

The Numbers (as of 2026-08-09)

Numbers that look like they are counting the same thing, but are not, turn up from page to page. They are only counting different things. All of them are correct.

NumberWhat it counts
about 6,000The hit count on my first search. Precisely, 6,037
5,728The number I actually read. It went down because an availability survey came partway through and some books dropped out of scope
883The number of notes I copied out as I read
879Of those, the ones with readable text. This is the number shown on the pages that read prefaces and afterwords
3,564The number used for the classification analysis. Counted from bibliographic data alone, regardless of whether the text is readable
1,580,026The count for the whole span of the Digital Collections, for comparison with the above
107The number of "statements involving quantities" gathered from personal records
365The number of daily behavior records the chronology is built on

The numbers grow. Notes are still being added, so the figures above are the ones from the day I wrote this, and they drift out of step with the numbers on each page.

The Four Biases

The numbers on this site carry the following four biases.

1. I am the one who chose.The notes are not a random sample. So the ratios between counts are not the proportions of public opinion at the time; they are the shape of one reader's subjectivity.

2. The bias of prefaces and afterwords.Prefaces and afterwords are written by the sort of person who writes and publishes a book. Anyone who could pull that off right after the defeat had energy to spare. That the vocabulary of construction outnumbers the vocabulary of collapse is, in a sense, only natural.

3. CensorshipGHQ's Press Code was fairly widely known. Which is to say, this writing was done with "censorship" in mind.

4. People who could publish a book were a minority.Right after the defeat, the people who could get a book out were a minority in society at large, and often in quite privileged circumstances.

Other Cautions

The text is raw OCR.Quotations are taken from the Digital Collections' text display, and misreadings I have not been able to fix are still in there. If you want the original, click through the link.

Currency figures are nominal values.5 yen in 1877 and 35,000 yen in 1966 cannot be compared as they stand.

Personal records are personal records.The "micro" side of the numbers pages was made by mechanically pulling everything containing a quantity out of the behavior records and the reading notes, and then sorting through it myself. The periods and the social strata are all over the place, and so are the sources. An article like the 1909 "eight-shaku serpent" is most likely the fantasy of whoever saw it.

Figures for the same item disagree with each other.The count changes with who did the tallying, and with the year and the place. On the numbers pages, things published between 1945 and 1949 are marked "contemporaneous figures" and things settled after that "later recounts."

The classification analysis has holes of its own.It leaves out the Prange Collection, and on top of that most kasutori magazines, cheap pamphlets, and underground publications fall out of it. Also, the "whole span" it is compared against includes fields that exploded in size from the high-growth era onward. Objective numbers, but a fine example of numbers with no objectivity in them.

What Is Not a Number

The tags I put on the prefaces and afterwords ("the mood of reconstruction," "hopes for the children," and so on, 35 kinds) were put there by me as I read; I did not decide the criteria for classification in advance, and I have not unified them after the fact. Notes with more than one tag are split at tally time and counted under both. So adding up the tag counts does not give you the total number of notes.

The tables where words were counted by morphological analysis were counted by machine. The list of which words to count I made myself, so that part is subjective.

Can This Be Reproduced?

The tallies come out of a script that reads the note files and counts them over again. The same file gives the same numbers. But I keep updating the notes, so counting on a different day gives different numbers. Which means every number on this site is a snapshot.