散文 · Gwern.net --- Essays · Gwern.net
|最后更新: 2025-12-25
Status
Like
Saved
Dec 25, 2025 02:48 AM
链接
notion image
This is the website of Gwern Branwen. I write about AI, psychology, & statistics. I am best known for my writings about AI scaling, poetry & anime neural networks; darknet markets and Bitcoin; blinded ; and dual n-back & spaced repetition.
Navigation: For information about my site’s philosophy & method, see the About page; for the website features & implementation, see the Design page; for information about myself, my use of other websites, and contact information, see the about-me page; for new pages, see the Changelog ( = new), short blog posts, or new links.
penAI announced in February 2019 in “Better Language Models and Their Implications”⁠ their creation of “GPT-2-1.5b”, a Transformer⁠⁠1⁠ neural network 10× larger than before⁠ trained (like a char-RNN with a predictive loss) by unsupervised learning on 40GB of high-quality text curated by Redditors. GPT-2-1.5b led to large improvements over GPT-1’s natural language generation, is close to or SOTA on natural language modeling⁠, and demonstrated high performance on untrained NLP tasks (see the paper for more details: “Language Models are Unsupervised Multitask Learners”⁠, Radford2019). By large improvements, one means that the best samples like the ones included in the OA announcement have started to reach an uncanny valley of text, capable of telling entire semi-coherent stories which can ⁠almost fool a sloppy reader—certainly, the verisimilitude is better than any char-RNN output I’ve seen. (A dump of many more samples is available on GitHub⁠. There is also an interactive word-by-word “GPT-2-Explorer”.) The full GPT-2-1.5b model was not released, but a much smaller one a tenth the size, GPT-2-117M was released in February 2019, which I call “GPT-2-117M” to avoid confusion.
GPT-2-117M was used in most initial experiments with GPT-2-based text generation. OA’s next largest models, GPT-2-355M & GPT-2-774M⁠ were released in May & August 2019, and the final, largest, GPT-2-1.5b model was released in November 2019⁠ (too late to be used in most of these experiments); 355M–774M turn out to just barely be trainable on commodity GPUs.⁠⁠2⁠ Also worth noting is the release of Gokaslan & Cohen2019’s independently-trained GPT-2-1.5b model⁠, which produces good samples if perhaps not quite as good as the OpenAI⁠ GPT-2-1.5b, but which was not trainable at the time on desktop GPUs⁠⁠3⁠ although it does still at least run (allowing for sampling/prompting if not training). OpenAI notes that “Since February, we’ve spoken with more than five groups who have replicated GPT-2”, and some have or gone further than GPT-2-1.5b: replications/extensions include Gokaslan & Cohen2019, GROVER⁠, Hugging Face (a NLP startup), XLNet⁠, Nvidia’s MegatronLM⁠, Google’s T5⁠ (finetuning Colab⁠), and MS’s DialoGPT⁠.
Naturally, people immediately used GPT-2-117M for all sorts of things, and I applied it myself to generate surreal anime plot summaries & dialogue for “This Waifu Does Not Exist”⁠.⁠⁠4⁠ Even more naturally, just as with char-RNNs, GPT-2 models, even unfinetuned, work well for poetry:
  • CTRL⁠ appears capable of generating verse when prompted with the “books” genre, see the Github repository’s “Weary with toil…” example⁠ (CTRL uses a ‘prefix’ approach similar to mine, and the “books” prefix corresponds to Project Gutenberg text, so it is not surprising that its samples would resemble my GPT-2-poetry samples)
Poetry is a natural fit for machine generation because we don’t necessarily expect it to make sense or have standard syntax/grammar/vocabulary, and because it is often as much about the sound as the sense. Humans may find even mediocre poetry quite hard to write, but machines are indefatigable⁠ and can generate many samples to select from, so the final results can be pretty decent.
The quality of the results is limited by sometimes only having access to smaller models and difficulty in running larger models at all; that can’t be fixed (yet). But quality is also reduced by GPT-2-117M being trained on all kinds of text, not just poetry, which means sampling may quickly diverge into prose (as seems to happen particularly easily if given only a single opening line, which presumably makes it hard for it to infer that it’s supposed to generate poetry rather than much more common prose), and it may not have learned poetry as well as it could have, as poetry presumably made up a minute fraction of its corpus (Redditors not being particularly fond of as unpopular a genre these days as poetry). Finetuning or retraining the released GPT-2-117M model on a large poetry corpus would solve the latter two problems.
The poetry samples above did not exploit finetuning because OpenAI did not provide any code to do so and declined to provide any when asked⁠. Fortunate, nshepperd wrote a simple finetuning training implementation⁠, which I could use for adding more interesting samples to my TWDNE and for retraining on poetry corpuses to compare with my previous char-RNN poetry attempts back in 2015–2016 (see the top of this page). An ⁠alternative GPT-2 training implementation⁠ with support for training on GCP TPUs has been created by Connor Leahy (technical details⁠), who trained a GPT-2-1.5b (albeit to substantially worse performance)⁠.
For the poetry corpus, Allison Parrish’s public domain⁠ “A Gutenberg Poetry Corpus”⁠ (“approximately three million lines of poetry extracted from hundreds of books from Project Gutenberg”) will serve admirably. A few other possibilities surface in Google Dataset Search⁠, like “Poems from poetryfoundation.org”⁠, but nothing particularly compelling.
As far as the text formatting goes, GPT-2-117M is flexible, you can dump in pretty much any text into a text file to use as the corpus, but some text formats are better than others. You want something which is as regular as possible (in both syntax & semantics), but also one which is as close to the kind of text you want generated, but also which wastes as few symbols as possible. Regularity makes learning easier, and you don’t want to have to massage the output too much, but on the other hand, GPT-2-117M has a narrow ‘window’ and no memory whatsoever, so if each line is padded out with a lot of formatting or even just whitespace, one would expect that to considerably damage output coherence—as most of the fixed ‘window’ is wasted on meaningless repetitive whitespace, while other changes like replacing newlines with the poetic convention of ’ / ’ are worse than nothing (since newline is 1 character vs 3 and maximally dense). Minimizing formatting also makes the cross-entropy⁠ loss easier to interpret or compare across datasets/runs: if there is a lot of formatting which is easy to predict or little formatting, the loss can look misleadingly good (or bad) as it easily predicts the formatting but struggles on more meaningful content.⁠⁠6⁠
The PG corpus has a strange format: each line is a separate JSON object, consisting of one line of poetry and a numeric ID for the work it’s from. Fortunately, the file as a whole is in order (if the lines were out of order, training on them would destroy the long-range language modeling which is the Transformer’s raison d’être!), so to turn it into a clean text file for training on, we can simply query it with jq and strip out the remaining formatting. This provides a pretty good format over all: the newlines are meaningful, no symbols are wasted on leading or trailing whitespace, and it looks like what we want. It is imperfect in that metadata/formatting we would like to be there, such as author or poem title, is not there, and things we would prefer not to be there, like the prose prefaces of books or annotations, are, but hard to see how to fix those easily. Another flaw I learned about only afterwards is that the PG corpus has been censored ad usum Delphini to remove⁠ “egregiously offensive content” such as “racist/sexist/ableist” words like “Pakistan” or “homogenous” (while, of course, permitting innocuous words like “s—t” or “f—k”). This corpus is unsuited for any serious academic work, but it should be fine for playing around with generating poems.
Setting up the GPT-2-117M training environment & obtaining the poetry corpus:
There is an additional step before beginning training. GPT-2-117M works with text in a “byte-pair encoding”, which is somewhere in between a character embedding & a word embedding. The point of this BPE encoding is that it is somewhat more efficient than raw characters, because it can chunk more common sub-words or phrases & this gets more complete words or phrases into the Transformer’s fixed ‘window’ of n symbols, but BPE still assigns symbols to individual letters, and thus arbitrary outputs can be generated, unlike word-level NNs which are more compact but trade this off by having a restricted vocabulary of m words seen in the training corpus and must treat everything else as the unknown token <UNK> (especially bad for rare words like proper names or variants of words like pluralization or tenses). The training code will encode the text corpus at startup if necessary, but for 117MB of text this is so slow that it is worth the extra work to run the encoding process in advance & store the results before training on it:
Yu-Gi-Oh meme: Kaiba vs Yugi. Kaiba attempts to trick Yugi into upgrading his Tensorflow installation, thereby risking breaking it permanently, as CUDA install problems are incomprehensible. Yugi upgrades, only for Kaiba to reveal his plan. However, it fails, as Yugi trains on CPU, accepting the extreme slowdown as the price for reliable software.
The temptation of CPU training after a ⁠bad Tensorflow upgrade⁠.
I assume you have a fully-working Nvidia CUDA & GPU-enabled Tensorflow installation, and have either run other DL code successfully or run the TF installation checklist’s MNIST⁠ toy example to verify that you have working GPU training. If you do not, I can give you no advice other than “good luck”. Debugging CUDA problems are the worst, and once you get a working setup, you should stick with it. If you just can’t solve the inscrutable crashes, you should look into using the free Google Colab GPU/TPU notebooks, or renting a cloud VM. (Training in the cloud is not as hard or complicated as it may seem, since you can pick a VM OS image which comes with CUDA/Tensorflow preinstalled and which will always work with the available GPUs.) You may be tempted to train on CPU to avoid the GPU-support mess entirely, but I advise against this: it will be at least 20× slower even if you have a many-core CPU like Threadripper, and it’ll save you time in the short run to switch to Colab or cloud or some alternative.
Then training proper can begin; my Nvidia1080ti⁠⁠7⁠ can fit a minibatch size of 2 (GPT-2-117M is still a large model), and I’d rather not see too much output so I reduce the frequency of checkpointing & random text generation:
The Python library “fire” used in the OA GPT-2 code is treacherous—it will not error out or even warn you⁠ if you typo a command-line option! Double or triple-check any new options you set against the available arguments defined by train_main in train.py, and keep this gotcha in mind if setting an option doesn’t appear to be doing anything.⁠⁠8⁠ While nshepperd has removed use of “fire” in favor of saner CLI options, watch out for this if you are using the original OA code or other derivatives.
Some hyperparameters could use tweaking:
  1. runtime, Temperature:
    1. ‘Temperature’ (0–∞) is used in sampling: the top-k most likely words are generated, and then selected randomly from; at 0, the most likely word is always chosen, while 1 means each is selected according to its likelihood, and it degenerates to a uniform 1 in k probability with higher values. In other words, the higher the temperature, the more chaotic or unlikely the generated sequences will be.
      In the original nshepperd code release, the default temperature setting for the samples during training, 1.0, is not the usual 0.7 everyone uses for GPT-2 prose sampling—although it turns out for poetry we don’t want it at 0.7 as that forces too many repeated lines & 0.9–1 turns out to be much better, so use temperature in that range when generating samples. (Higher still may be better but I have not experimented with >1.)
      If you are sampling after 2019-05-15, it may be a better idea to use a new sampling strategy, “nucleus sampling”⁠ (which essentially sets a different k at each step to avoid sampling extremely unlikely words and greatly reduces the repetition problem), which can be enabled like --top_p 0.9. (An interesting but untested sampling strategy is ⁠“tail free sampling”.)
  1. train time, Learning Rate (LR):
    1. A key NN hyperparameter as always.
      In nshepperd’s code, the Adam SGD⁠ learning rate is left at its TensorFlow default of 0.001, which works initially, but appears to be much too high for this purpose (perhaps because the minibatch is so tiny on 1 GPU). After training overnight, the loss was not decreasing below 2.5, so I decayed it manually to 0.0001 & resumed training (editing line 136 of train.py to read tf.train.AdamOptimizer(learning_rate = 0.001*0.10)), eventually decaying it again (to 0.001*0.0001) to get it down to a loss of ~1.95. (nshepperd has since added a --learning_rate option so manual editing of the source is no longer necessary.)
After training GPT-2-117M an hour or two, a sample
Overnight samples during training:
The loss here is the usual cross-entropy we often see in architectures like a char-RNN. Typically, the best text generation results come when the model has trained down to a cross-entropy of <1, and 2–4 tend to be incoherent gibberish. (For example, in Andrej Karpathy’s Tiny Shakespeare.) That loss is per character, while GPT-2 operates on BPEs, which usually encode multiple characters, so are harder to predict; it seems to me that the conversion factor is ~2–3, so a GPT-2 model should aim for a loss of <2 if a good char-RNN would reach losses like <1. In this case, GPT-2-117M’s original poetry modeling capability is not too shabby (as demonstrated by the various prompted samples), and it shows decent poetry samples starting ~3.5. (Gibberish seems to set in at losses >6.) Given how large & powerful GPT-2-117M is, even with this much poetry to work with, overfitting remains a concern—memorizing poetry is not amusing, we want creative extrapolation or mashups.
For this model & dataset, I trained for 519,407 steps to a final loss of ~2 in 72 GPU-hours; almost all of the learning was achieved in the first ~16 GPU-hours, and training it additional days did not do any apparent good in terms of the loss itself.⁠⁠9⁠ This suggests that GPT-2-poetry was underfitting the poetry corpus & would benefit from an even larger model size.
Downloads:
Before sampling from any new finetuned version of GPT-2-117M, remember to copy encoder.json/hparams.json/vocab.bpe from the GPT-2-117MB model directory into the new model’s directory. I find higher temperature settings work better for poetry (perhaps because poetry is inherently more repetitive than prose), and top-k appears to work fine at OA’s top-40. So unconditional sampling can be done like this to generate 2 samples:
Not bad.
hawn Presser⁠ noted the issues with the Project Gutenberg corpus and, as book-level transitions are solved, suggested a heuristic for reconstructing the blank lines denoting (presumably) stanzas: use the numbers (GIDs) from the original JSON for book-level transitions, and look for lines which might be transitions to insert newlines. (Imperfect, but better than nothing!) Since stanzas are still connected, <|endoftext|> is used for the book-level transitions, and a blank line is used for the stanza-level, preserving as much of the semantics as possible.
This is converted to the NPZ and trained as usual. I retrained the previous non-prefix GPT-2-117M PG poetry model for ~30k steps (>16h?) down to a loss of ~1.73. (I used GPT-2-117M instead of GPT-2-345M for compatibility with my concurrent experiment in preference-learning training.)
The results are quite nice, and competitive even with ⁠345M⁠. Some selected samples:
While training that, I recalled that my other major beef with the PG corpus was its absence of more contemporary poetry. I didn’t really feel like trying to scrape the Poetry Foundation (PF) website myself, but I gave Google Dataset Search another try and to my surprise, discovered a scrape had surfaced on Kaggle⁠. Aside from being large, it comes with interesting metadata: the title and author, but also a somewhat eccentric set of ‘tags’ describing the poem. They would be nice to use via the ⁠inline metadata trick⁠, allowing some degree of controllability (like my use of author or book ID prefixes, or CTRL’s use of subreddits).
The Kaggle Poetry Foundation scrape has numerous issues I had to fix before its format was acceptable for GPT-2:
  • I replaced prefixed whitespace, trimming leading/trailing whitespace in all fields
  • replace 3+ spaces with newlines
  • deleted all 2+ spaces
  • dropped poems with <100 characters (generally a scrape error)
  • remove Unicode junk
  • serialize it as title+author+tags (if any) / poem /<|endoftext|> (ie. the inline metadata trick, allowing for potentially better learning and a small degree of control in conditional generation)
Once the cleaned PG was done, I then swapped out the PG for PF dataset and began finetuning. (I could train on the combined dataset, but the PF dataset is only 20MB and at 1⁄6 the size of PG, training on that would take a long time to pick up on PF’s details.) Surprisingly, the PF dataset trained down to ~0.60 loss after ~10k steps, as compared to PG’s ~1.73, a decrease much greater than I had expected from providing some metadata, suggesting that modern poetry, being so prose-like, is much easier for GPT-2 to model—which doesn’t strike me as a compliment.
The contemporary PF samples properly mimic all the metadata and formatting, and are good, for what they are. (If you doubt this, read through a random selection of PF poems.) The ones I liked seemed like they benefited greatly from the PG pretraining. There are still a lot of flaws in the unclean PF data: run-on lines are particularly irritating, and appear to be flaws in the Kaggle scrape dataset rather than the original PF website. I have brought up the problems on Kaggle, but I doubt they’ll be fixed soon.
With PF done, I combined it with PG and trained on the combination dataset for another ~20,000 steps, yielding a final loss of 1.61. The combined model appears able to do both datasets well (the weighted average of a dataset with a loss of 0.6 and another dataset 6 times larger and a loss of 1.7 would be ~1.55, close to the model’s ~1.6), and the samples don’t appear to differ much, so I don’t excerpt any. But the combined model should make a great starting point for RL preference-learning training.
The first version of the PG data for GPT-2-poetry just runs all the lines together, erasing the metadata about what book each line comes from. A good model should nevertheless gradually learn about the transitions between poems & whole books, but that is hard and there may not be enough transitions in the data to learn effectively.
Much like the char-RNN experiments on this page, there is no reason one can’t inject that metadata in a structured way to see if the model can learn to exploit the metadata; even if it cannot, the added metadata shouldn’t hurt that much because it is so regular & repetitive. Inserting the metadata also allows for some degree of control in conditional generation; one should be able to put in the book ID for, say, Homer’s Iliad as a prompt and get out a long block of consistent Homeric pastiche.⁠⁠10⁠
Ideally, there would be unique IDs for every author, poem, and book and these would appear at the beginning of every poem and the end of the poem would be delimited with the <|endoftext|> symbol that OA’s GPT-2 models were trained with, but unfortunately only the book ID is available in this particular dataset. (Project Gutenberg ebooks do not include any metadata or formatting which would cleanly split each discrete poem from each other.) Like before with authors, the book ID metadata can be formatted as a prefix on every line with a delimiter like the pipe character.
Rather than start over with GPT-2-117M again, GPT-2-poetry can just be further finetuned on this new prefixed version of the PG corpus to produce what I call “GPT-2-poetry-prefix”:
The loss of GPT-2-poetry-prefix will be much lower than GPT-2-poetry because the prefix is so predictable, but it will hopefully learn interesting things beyond that.
In other samples, the generated IDs switch in the first two lines, and while that’s not much to judge from, GPT-2-poetry-prefix seems to ignore keywords from the first line when the IDs change, and doesn’t repeat them in the rest of the sample or attempt to rhyme off them, which is further evidence it is successfully associating & learning to mode-switch.
Like GPT-2-poetry, GPT-2-poetry-prefix converged quickly to a final loss of ~1.6 after 224,474 steps taking 31 GPU-hours, not improving much after the first ~8 GPU-hours despite decreasing the learning rate. (Diminishing returns⁠ appear to set in quickly for finetuning GPT-2-117M even if one has a relatively large new corpus.)
One training sample is worth remarking on:
The rhyming in this sample is so good as to be suspicious. It might also sound familiar—because many of these lines are being copied from Thomas Gray’s⁠ Elegy Written in a Country Churchyard⁠, which opens:
Some spelling differences aside, this intro is almost entirely copied from the 8 copies of Gray’s poem in the corpus; this extensive copying is not something I spotted in the GPT-2-poetry samples I looked at, suggesting that the scaffolding of the metadata did indeed help with learning.
Also interestingly, the copying only goes so far, as immediately after the final line about the owl, where Gray continues:
GPT-2-poetry-prefix instead continues:
That is, it focuses on the female figure of the Moon in a way more ode-like than elegiac.
These lines also do not seem to be extracted from the rest of Elegy either, as words like “bliss” or “mirage” or “dream” or “seraph” or “Platonic” do not appear in it. Some of the phrases like “blissful dreams” do appear in the rest of the corpus, but others like “some mirage” or “mirage she” do not. Nevertheless, the style is consistent throughout the entire sample and the quality is good, suggesting that while GPT-2-poetry-prefix has managed to memorize to a limited extent, it is nevertheless fully capable of generating good original text.
An additional example of memorization has been spotted; sample #17 in the 1,000 unconditional samples is almost entirely a memorized copy of Percy Bysshe Shelley’s ⁠“To a Skylark”⁠:
The 87 lines beginning with “Hail to thee, blithe Spirit!” are all Shelley (with perhaps slight spelling differences), much surpassing the memorization for Thomas Gray. Considering the top-k sampling method, it’s amazing that the sample could so exactly follow “To A Skylark”. It turns out there are ~12 copies of the poem in the PG corpus (it’s a popular poem), so in retrospect some degree of memorization is not surprising, but that’s still a lot of memorization. The 4 lines beforehand don’t appear to be copied from another Shelley poem, making it even more amazing. It’s a pity that that sample did not continue further because one wonders whether it could have repeated the entire poem and what it would’ve done when the original poem ended.
For both GPT-2s, I generated 1000 samples as follows:
Download links again:
Some fun passages I noticed in the first 100 unconditional samples:
Not bad.
These samples represent roughly top decile poem samples (~10 out of the first 100), at least by my selection.
Scott Alexander highlights a fun repetition-trap one:
The top percentile of poems are probably quite good, especially with some light human editing to fix up the more glaring issues. To get a decent number of top percentile poems would require a lot of reading, but on the other hand, there is no reason why selecting or ranking poem samples could not itself be treated as a supervised learning task for retraining GPT-2-117M-poetry on, by using selected/non-selected as labels and training to predict the probability of a given sample being selected, and then such a NN could be used to prioritize likely-good GPT-2-poetry poems (or any source of poetry) for human review (and, in a form of “active learning”, the results of the manual review can be fed back in as additional data to help discriminate between the best and the merely good samples).
Prompted samples can be done like this:
The downside of using the stock OA interactive prompt is that it returns on the first newline, so one either deletes newlines or uses a single line. Neither is good: a single line is hardly any context, while smashing many lines into a single super-long-line is dangerous because neither GPT-2 has ever seen poems formatted that way (only, perhaps, some prose that snuck in) and newlines have important semantic functions in poetry. So, to avoid either problem, I bypassed the interactive prompt entirely, and I modified the Python script to replace input (for taking 1 line of keyboard input) to instead read standard input (import sys; sys.stdin.read()) so I could simply pipe in multiple lines from files or from the copy-paste buffer using xclip -o.
The next issue in prompts is the question of the metadata; given that all the training data was properly labeled with origin and learning the meaning/associations was much of the point, it doesn’t make sense to not exploit this control in generation. If I was using authors, as with my previous char-RNN experiments, the prefix is simply whatever author one wants completions from, but in this case, it’s not quite so simple since we only have book IDs
If an author is already represented in the PG corpus, hypothetically one could look them up in it and see what IDs their poems were included under and use that, but that is a pain and doesn’t work for ones outside the corpus like Ginsberg. So, one could instead simply ask the model what prefix it thinks a prompt should use by feeding in the input several times and seeing what prefix it confabulates⁠ in the samples, and then adding that to the input for the real samples. If GPT-2-poetry-prefix consistently returns a specific prefix, then that is what it has learned and is useful scaffolding for the inputs; if it can’t do so consistently, then the prefixes aren’t useful for this particular input and it doesn’t matter.
So, to generate samples conditioned on relevant metadata, I pipe in the whole input unmodified several times, look at the generated samples for an ID, and if there is a consistent ID, then prefix it to the input and sample again several times.
Of course, now that everything is trained & I have a good input method, I want to see how GPT-2-poetry-prefix does on the same poems as GPT-2-117M before!
First, “Howl”. Given that the Project Gutenberg corpus is entirely old poetry and wouldn’t include much in the vein of “Howl”, I didn’t expect this to be good. The finetuning would wipe out the knowledge of free verse.
Finding a good prefix was hard, also unsurprising—not much like it in the PG corpus! I ultimately had to settle for a “1997” prefix from a relatively free-verse sample for generating the 3 samples:
While they may be OK on their own and plausible as unconditional samples, they are disappointing as conditional completions, largely ignoring both the vocabulary & style. It would seem that the finetuning wiped out whatever it was GPT-2-117M was using to generate its amusing “Howl” completions.
For “Ozymandias”, I fed it in a few times, and it seemed to like numeric IDs starting with ‘88’, so I used this as a prompt:
Yielding (3 samples):
Sample #2 is over-influenced by some prose footnotes/commentary which apparently were in PG, but the analogy of Ozymandias to Aztecs is a potentially fruitful one. And sample #3 here is a particularly good extension.
Not clear what text exactly ⁠Scott Alexander⁠ used from Alexander Pope’s Essay, so I quoted the famous beginning section of Part 2. 3 samples strongly indicated Pope-like writing was associated with a prefix of ‘385’ (if not necessarily a full prefix) so I used 38511 for the following 3 samples:
Alexander described his GPT-2-117M sample from Pope:
GPT-2-poetry-prefix still has “overeducated 18th-century dandy” down pat, but it manages to improve on the rhyming aspect: there’s quite a few rhyming lines in samples #2 & #3 (#2 seems to be screwed up by taking a digression into footnotes defining words and then bad sampling getting it trapped), like “pretence”/“sense”, “soul”/“whole”, “love”/“glove”, “state”/“Fate”, “bright”/“sight”, and a number of almost rhymes like “right”/“great”. One wonders if it’s learning by brute force and memorizing specific pairs of rhymes (although could there really be that many rhymes of “state”/“Fate” in even 3m lines of old poetry?), or if it’s doing something more equivalent to inferring the latent⁠ phonetics from the co-occurrence of byte-pairs? (That may sound unlikely, but word embeddings do many unlikely-sounding things with no more supervision than co-occurrence⁠⁠11⁠.)
More concerningly, the samples are terrible. Pope’s poetry should be straightforward for GPT-2-poetry-prefix, as it follows standard meters and rhyme and relies on a classical vocabulary well-represented in the PG corpus.
Why, then, are they so bad? I suspect this may reflect the corpus itself doing Pope a disservice. Pope’s inclusion in the ⁠PG corpus appears to consist of the following (grepping for “Alexander Pope”):
Checking PG entries and looking through the 32190 prefix, it starts:
This is perhaps not good training material for GPT-2-117M-poetry/prefix and explains the bizarre degeneration—it is ‘expecting’ sudden random irruptions of largely-irrelevant prose such as introductions or footnote-annotations (rendered inline by PG’s text formatting).
Other entries in the corpus will be more free of scholarly or prose apparatus. (In retrospect, a preprocessing step like dropping lines longer than ~60 characters might’ve been a good idea.)
The prefix trick doesn’t work on the 8 famous first lines nearly as well as it does with the long excerpts from “Howl” etc; I assume they are simply too short to home in on a relevant prefix. Nevertheless, I tried.
“It little profits that an idle king,” yielded no consistency in prefixes, so I skipped adding one. 3 samples:
“That is no country for old men.”, no consensus. 3 samples:
“Come, my tan-faced children,”⁠; no consensus, 3 samples:
“Let us go then, you and I,”; no consensus, 3 samples:
“To be, or not to be: that is the question:”; some consistency, so prefix “1006”; 3 samples:
“Romeo, Romeo! Wherefore art thou Romeo?”; some consistency, with 1006 popping up again as a prefix (Shakespeare perhaps is memorable enough for GPT-2-poetry-prefix); 3 samples:
Upon request, I ⁠generated 100 samples⁠ of Lewis Carroll’s “Jabberwocky⁠”. Examining preliminary samples, the closest prefix was #24650, corresponding to ⁠The Jingle Book⁠, Wells1899, an anthology of humorous children’s verse (which makes sense). “Jabberwocky” itself does not appear in the PG corpus but the “Jabberwock” is mentioned in one of the poems in Wells1899, the acrostic poem “An Alphabet Zoo”, so, close enough.
In May 2019, OpenAI released the next-largest model⁠, which increases the parameter count from 117 million to 335 million, an increase of almost 3×. The GPT-2-345M model has increased layer depth & more attention heads but apparently similar window size; as such, while it may not be much more able to maintain coherency across long samples, its coherency & quality should be superior to GPT-2-117M within each window, as it can absorb more knowledge into its parameters & the increased depth may allow for more ‘thinking’ at each step.
The regular text samples from the GPT-2-345M model struck me as somewhat subtly but noticeably higher-quality than GPT-2-117M, so while I was hoping someone would supersede GPT-2 entirely by releasing a more advanced model (like a large-scale Transformer XL or Universal Transformer⁠, or even newer models like the UniLM⁠ which marries bidirectional & unidirectional Transformers), I decided to train GPT-2-345M on the PG corpus to compare it with GPT-2-117M.
This proved more difficult than GPT-2-117M. The GPT-2-117M model was already large, at 480MB for the whole, so making it 3× larger bloats it to 1.4GB on disk; and the VRAM use on a GPU is even worse: with GPT-2-117M, a training minibatch of n = 2 could barely fit on a 1080ti’s 11GB, but at GPT-2-345M, n < 1! The main culprit seems to be the self-attention layers, as regular self-attention scales more than linearly, so GPU VRAM gets eaten up fast, and apparently 16GB might not have been enough for GPT-2-345M either. While I have enough system RAM to train GPT-2-345M without any tricks, my Threadripper CPU is still ~14× slower than a 1080ti, and if one guesses that GPT-2-345M takes 3× longer to train than GPT-2-117M, and GPT-2-117M takes 1–2 days, and CPU is 14× slower, then that’s <84 days for the poetry finetuning, which would not be fun.
To solve this, nshepperd extended his GPT-2 training codebase to employ a technique OpenAI helped introduce (and presumably used in training GPT-2, although the GPT-2 paper is silent on the details): “gradient checkpointing”⁠. Gradient checkpointing is a space-time tradeoff which throws away some of the intermediate states of a NN, potentially greatly reducing total VRAM use, but at the cost of some slowdown when those intermediate states need to be recomputed for doing the backpropagation⁠; the slowdown, fortunately, turns out to be fairly modest.
The downside of gradient checkpointing is that for GPT-2-345M, it is still not memory-efficient enough to train it just like GPT-2-117M—the self-attention layers checkpoint nicely (as the Sparse Transformers paper⁠ remarks⁠⁠12⁠ apropos of needing extremely wide Transformer windows to accomplish MuseNet⁠), but it’s not enough, due to the giant word/BPE-embedding, which blows out RAM usage. (Although it’s possible nshepperd didn’t implement gradient checkpointing quite right for GPT-2, as the OpenAI papers don’t mention any difficulties related to the embedding or using gradient checkpointing.) His initial solution was to simply disable training of the embedding and train only the Transformer layers, reasoning that the generic English embedding probably wouldn’t need to be trained that much as the Transformer layers are where the real work is done; much later, it occurred to us that the Adam SGD optimizer was part of the memory problem, as, being an adaptive ⁠momentum⁠-based SGD optimizer, it must store a mean/variance⁠ for every parameter to adjust its updates per-parameter, which greatly increases memory use (and which gradient checkpointing does nothing about); when we switched to simple SGD, that freed up enough RAM to re-enable the embedding. This is important in part because the learning rate for Adam & SGD differs by orders of magnitude: a LR <0.01 seemed good for me for SGD, but Adam wanted a LR more like 0.00001. With n = 1 minibatches, the training loss is extremely noisy and it is difficult to see the impact of any hyperparameter changes for the usual hand-tuning, so nshepperd also implemented a simple ‘validation loss’ function, which was helpful toward the end.
So, the upshot seems to be that GPT-2-117M can be trained end-to-end⁠ with Adam on a commodity GPU in 11GB VRAM; and GPT-2-345M must be trained with gradient checkpointing, and one must choose between either fancy SGD optimizers or full end-to-end training including the embedding; and 744M (released 2019-08-20) can’t be trained at all. Toward the end, I switched from Adam+Transformer-only to SGD+all, and this seemed to drop my GPT-2-345M-poetry validation loss by ~0.01 to a final 1.915 (which is not nothing, so perhaps the embedding did need some adjusting for a more poetic vocabulary).
In total, I trained GPT-2-345M-poetry for 815,326 steps (minibatch n = 1), with an Adam LR ~ 0.00001 and SGD LR ~ 0.001, over ~7 days (2019-05-04–2019-05-13) on 1 Nvidia1080ti; the necessary training time, with the benefit of hindsight, was probably closer to 3 wallclock days. GPT-2-345M-poetry converged to a final loss of 1.915, an improvement of ~0.1 over GPT-2-117M’s ~2 loss (so, in some objective sense, which is indirectly related to generated poetry quality, one could say that GPT-2-345M is 5% better than GPT-2-117M). I had expected somewhat more quantitatively, so I wonder if more aggressive training methods like cyclic⁠ learning rates⁠+SWA⁠ would have worked if they were implemented in this codebase & I had the patience to wait a week or two for multiple cycles? In any case:
Transcript: [Megan is sitting at a computer, and Cueball is standing behind her.] / Megan: ‘Looks like computers will beat humans at Go pretty soon.’ / Cueball: ‘Wow.’ / Cueball: ‘That’s the last of the big ones.’ / Megan: ‘Yeah.’ / [Megan looks back over her shoulder at him.] / Cueball: ‘Well, at least humans are still better at, uh...’ / Cueball: ‘coming up with reassuring parables about things humans are better at?’ / Megan: ‘Hmm.’ / [Megan types on her computer.] / *type type* / [She leans back over her chair again and addresses Cueball.] / Megan: ‘I made a Python script that generates thousands of reassuring parables per second.’ / Cueball: ‘Dammit.’ / Computer: ‘Computers will never understand a sonnet computers will never enjoy a salad comp—’
Testing GPT-2-345M-poetry, a slightly higher temperature felt warranted, so to generate ⁠5000 random poetry samples⁠:
I also generated ⁠500 conditional samples⁠ for Yeats’s “The Second Coming”.
Reading through training & random samples, they feel noticeably more coherent; it feels easier to extract meaningful subsections which form reasonable poems. (In particular, the pastiches of classical epics or Dante have gotten remarkably good.)
Here is a ‘failed’ example, where GPT-2-345M-poetry imitates the scholarly apparatus that unfortunately contaminates the PG poetry corpus; it is quite plausible-sounding, even including plausible-looking Latin:
Ganbare, GPT-2-chan:
This is a peculiar one; it starts as a satirical poem but I can’t make out what it is trying to switch to partway:
This one I think must be a mix of The Song of Hiawatha⁠ and the Kalevala⁠ (but if a wizard offers you rainbow-colorful draughts of rum strained through his magic red beard, I suggest declining in the interests of hygiene):
An alt-history where Germany won WWI:
The Tao Te Ching⁠ (TTC) is a famously enigmatic text, written in a difficult style in an more difficult language, and because of the challenge, has attracted many highly-varied translations. (For a 2024 attempt using ⁠ChatGPT⁠, see ⁠“a wandering mind”.)
⁠Hiræth Weltschmerz⁠ compiled ⁠~270 translations of the first verse of the TTC⁠ in one text (108Kb/19k words), and compiled per-chapter translations of the rest of the TTC as well (4.8M/870k words). Using TTC snippets to prompt GPT-2-345M didn’t produce good results, so he asked me to train a GPT-2-345M on the TTC to see what it did.
I converted the first corpus to Unix⁠ text format, replaced various escaped character entities with their Unicode equivalents, and replaced double newlines with single lines, and trained the original GPT-2-345M (I was unsure if using one of the poetry models would help) with the usual settings for ~6 GPU-hours, at which point it had reached a loss of ~1.8 and I began to worry about overfitting & stopped.
I generated the usual ~1k unconditional samples:
Some training samples:
I then began training it on the full TTC corpus, which was split into per-chapter files. Remembering the problems with run-on poetry, I added the <|endoftext|> markers to the end of each file. (It would be better to add that to the end of each translation, but HW didn’t include any delimiters for each translation, and doing so manually would be too much work.)
Hiræth Weltschmerz was able to improve the TTC training dataset by providing the first 41 chapters with the original newlines/linebreaks, and separate translations separated by a blank line, so I replaced blank lines with <|endoftext|> and trained that for ~24 GPU-hours to a final loss of ~2.10. The results read much more poetically, I felt.
For this final TTC-with-linebreaks, I uploaded the model & generated 1,000 random samples as usual, but I also generated per-chapter samples. For per-chapter samples, I used csplit to split each file/chapter into separate verses/translations, selected a random 10 of them, and then generated 10 random completions of each one (100 total per chapter):
The idea there is that one can write one’s own Tao of GPT-2, going chapter by chapter: select some of the chapter 1 prompted conditional sample completions to create a new chapter 1, and so on, in a way which would be difficult to do with just random unconditional samples.
In keeping with its gradual rollout plan, observing no particular misuse in the wild (aside from a few anecdotes about content mills), OpenAI released the final and largest model, GPT-2-1.5b, in November 2019 along with detection tools⁠ (paper⁠) The model was an easy upgrade for services like Talk To Transformer which simply sample from the original model, since it still fits easily onto commodity GPUs.
From November–December 2019, ⁠Shawn Presser⁠ & I worked on finetuning training GPT-2-1.5b on the combined PG+PF poetry dataset from above. The 1080ti GPU approach failed, so we switched to Google Colab to use the free TPUs. Colab worked, but constant failures made it painful to contemplate multi-week training runs, and so we switched to GCP to use TPUs directly. Direct TPU use is much faster, but the errors remained, so we began working on a distributed TPU approach, to work around individual TPU errors. Eventually, using Google TRC research credits to pay for TPUS, we began running ⁠TPU ‘swarms’ of <60 TPUs (since scaled to <200). These produced meaningful training progress and we reached a loss of ~2 by 2019-12-12.
We were able to train it to ~1 loss, but it appeared to have overfit in some fashion as sampled qualities became increasingly worse by the time we halted ~2019-12-20, so we settled for iteration #500,522.
Training is a different story. 345M took ~7 days to train, and GPT-2-1.5b is 4.4× larger, so that alone implies a training time of a month. Worse, where GPT-2-345M fits in reasonably in a 1080ti’s 11GB VRAM & 745M just barely fits, GPT-2-1.5b does not fit at all.
We began trying to train on the combined PG+PF corpus, like the GPT-2-117M model I trained for RL preference learning⁠, but turning on all the options in the nshepperd repo doesn’t fix the memory problems. (FeepingCreature was able to train on his new AMD⁠ GPU, which has 16GB RAM, so a few more gigabytes would’ve done the trick..)
Using Shawn Presser’s fork of nshepperd’s fork⁠, we experimented with alternatives like using reduced-precision, truncating parameters to FP16. This caused serious errors.⁠⁠13⁠ After fixing those errors, and reducing the context window by half (potentially hamstringing it), we could train GPT-2-1.5b on a 1080ti, but our naive conversion to FP16 appears to have seriously damaged the model and it emitted only garbage. We then tried using a different floating-point format, bfloat16⁠, which in theory is much better suited to NN models than FP16 & natively supported on TPUs, but it trained extremely slowly on my Nvidia1080ti GPU. Given the daunting expected training time, bfloat16 was not a solution.
The only solution here seemed to be to abandon my 1080tis and upgrade to TPUs. TPUs may not be any faster, but they have far more RAM and can train a GPT-2-1.5b with no problem.
How to get a TPU? Fortunately, Google Colab did just enable free TPUs by default… So Presser enhanced his fork to support TPUs, and we started training.
Unfortunately, Colab notebooks are still limited in system RAM and disk space, so training GPT-2-1.5b then encountered the surprising problem of running out of RAM & crashing, running out of disk space, and saving to disk being extremely slow due to slow TensorFlow serialization of the model checkpoint. (The TPU-based serialization code would have been far faster using the standard TF way, but it would also required the user to create & manage a Google Cloud bucket; we were still hoping to create an easy works-out-of-the-box Colab notebook to let anyone do GPT-2-1.5b-finetuning. If there was a faster way to do it, Presser didn’t know about it.) This was partially solved by saving few checkpoints, figuring out how to attach a Google Drive folder⁠⁠14⁠ (after paying $2.54/month for an upgrade to ~100GB of additional space, since the default 15GB Google Drive is perilously small), and further work on optimizing the serializing. Training was slow—1 minute per minibatch, initially—but did work. An example:
Presser tried out curriculum learning/progressive growing by setting the context window to a small window like k = 50 BPE tokens, with the idea that it could be gradually annealed to the original k = 1,024 over the course of training. (Because of how Transformers scale, k = 50 uses far less memory & compute than k = 1024, so it fits much larger faster minibatches.) This seemed to be working to some degree, but it was no silver bullet.
Exacerbating the problem, TPUs on Colab appear to randomly ‘freeze’, an issue unrelated to Colab notebooks timing out after a day or so; manually interrupting the training process and restarting fixes it, but at the cost of any progress made since the last (slow) checkpoint & required constant babysitting; I calculated that one would have to checkpoint every hour to optimize the tradeoff between freezes & checkpoints! At one point I was dealing with TPU freezes every half hour. It was already giving decent poetry samples despite a loss >3, but we wanted to train to convergence, which ought to be <1.6 (the final combined-117M loss).
This was not going to work for weeks of training. Presser again modified the codebase & notebook to add ‘watchdog’⁠ processes which would watch for an apparently hung TensorFlow process due to a TPU freeze, and kill it and restart. But the lost time was a serious issue: we couldn’t checkpoint too often because then we’d waste all our time checkpointing, but not checkpointing meant we’d lose minutes or hours of training. We couldn’t find any information about why TPUs would freeze and figured it was some sort of Colab issue, so I decided to bite the bullet and pay for a GCP VM & TPU.
⁠A preemptible TPUv2⁠ costs $1.72/hour, which is not too bad… A week of training would cost >$287 & I’d never used GCP before, but I was too curious what a fully-trained GPT-2-1.5b would generate. The net cost for November 2019, due to all the experiments and costs not covered by the TRC research credits, was $409, primarily for high-RAM instances, and then network egress bandwidth fees—for cross-zone traffic with the TPUs, apparently.⁠⁠15⁠ Optimizing for GPT-2-1.5b-poetry, we got December 2019’s cost down to $254, and trained GPT-2-1.5b-poetry and an IRC logs model. We spent in January 2020 an additional $512 on a number of projects: the chess⁠, Subreddit Simulator⁠, Archive Of Our Own, and video game walkthrough GPT-2-1.5 models; the 30k context window GPT-2-117M ABC/MIDI model⁠; ImageNet resnet⁠ benchmarking; and StyleGAN⁠ 2 prototyping for training on ⁠Danbooru2019⁠.
After setting up on GCP and figuring out the details like needing to set an environment variable with the target TPU name, we discovered… the TPUs kept freezing anyway. This was on top of the standard preempting of TPUs, since we were using preemptibles to save money as is standard in cloud deep learning.⁠⁠16⁠ This remained a mystery. Was there some undocumented heartbeat⁠ that was required? Was Presser’s TF code, which avoided the standard TPUEstimator approach which it seems everyone else uses, triggering some sort of problem? (We were warned in vague terms that TPUs do not like loops or reshaping operations.) Even more irritatingly, our on-demand TPUs turned out to preempt anyway! But at least the checkpoints were fast, so now the watchdogs worked better. But on the gripping hand, the TPU was not fast and was performing far below what we thought it should based on its nominal specs, and we seemed to be using barely a third of the cores.
GPT-2 model training curves over 4 days on 1 TPU each: GPT-2-1.5b (orange), 774M (light blue), 400M (red), ‘tiny’ reduced context (dark blue)
GPT-2 model training curves over 4 days on 1 TPU each: GPT-2-1.5b (orange), 774M (light blue), 400M (red), ‘tiny’ reduced context (dark blue)
Presser decided to press on and after further optimizing work to ensure we used the full TPU RAM and more of the cores⁠⁠17⁠ with a minibatch n = 4, began experimenting with support for multiple TPUs. Since each TPU is a separate computer inside Google’s network, and not ‘attached’ to a VM like a GPU, there was in theory little limit to how many TPUs our 1 VM could orchestrate. We could, in theory, create an equivalent of the expensive TPU ‘pods’ by simply connecting to a bunch of TPUs at once.
The main limit for distributed TPU training is the network bandwidth: copying around the latest version of the multi-gigabyte model to and from the central VM uses up all the bandwidth available. In the simple synchronous case, which most closely approximates training on a single GPU, if the entire cluster has to stop and wait for every node to copy its updates to the master, the master do a single batch update, and the master sync back out to each node, the cluster will spend most of its time just waiting on the network to copy everything. (And what happens when one or more TPUs inevitably freeze?)
Presser worked around the bandwidth with an asynchronous approach somewhat like the old HOGWILD⁠ training method: instead of every node copying its entire model at a fixed timestep and waiting for all the other nodes, the nodes are constantly communicating a fraction of their latest model with the master and receiving an updated fraction back, regardless of how many iterations other nodes have run. So every node is running a hybrid & partially-out-of-date model and sending stale gradients out, but gradient descent is robust enough that this will still work and will scale up easily. After enough slices have been sent, a node will have sent an equivalent of a full model, and caught up partially, and the ‘swarm’ will hopefully be able to make progress by training on a large amount of hardware and be faster than using just a few TPUs synchronously.
A swarm was too expensive for me, so we applied for TensorFlow Research Cloud credits⁠. I wasn’t expecting anything to come of it, but the form was easily to fill out in a minute (it’s not much more than an email address), and to my surprise, within 3 hours we had been approved for 1 month of credits, covering several on-demand TPUv2–3s, and 100 preemptible TPUv2s⁠ (but no TPU pods).
Presser began the long and painful process of debugging the swarm and all its problems… The halts were never quite fixed but we kept scaling.
By 2019-12-11, after applying for additional credits because we were coming up on the TRC deadline on the 14th, we’d gotten the loss down to ~2.15. After switching to Adam and scaling the swarm further to ~95 TPUs on the 13th, we reached a loss of 1.61, matching or beating the GPT-2-117M record on the combined dataset. A further 5 days (interrupted by swarm preemption and occasional tweaks/experiments) brought the loss down to <0.6 on 2019-12-18. I had expected flagrant plagiarism/overfitting well before a loss of 0.6, perhaps ~1.2, but regularly inspecting unconditional samples and searching initial lines or generated titles/authors, I found little & they didn’t read like plagiarism, so we kept training to see how far it could go. (Prompting with lines from famous poems would’ve almost surely elicited plagiarism, but I am less concerned with that, since GPT-2-1.5b is so big it can easily memorize famous poems without compromising its general poetry abilities.) I suspect that GPT-2-1.5b is not really >3× better than GPT-2-117M, and that GPT-2-117M could have been trained to <1.6 loss if we had used similar amounts of compute, so the actual benefit from scaling up GPT-2 is smaller—but why bother with training GPT-2-117M to a better convergence when we can use GPT-2-1.5b?
Example training curves:
Training curve of a swarm of ~97 TPUs training GPT-2-1.5b-poetry for ~21 hours (2019-12-13) from a loss of ~2.15 to <1.6.
Training curve of a swarm of ~97 TPUs training GPT-2-1.5b-poetry for ~21 hours (2019-12-13) from a loss of ~2.15 to <1.6.
100 TPUs, 1.6 → 1 loss (2019-12-16)
100 TPUs, 1.6 → 1 loss (2019-12-16)
While sampling, we noticed double and single quotes were being replaced by mojibake⁠ gibberish. This appeared to be due to Unicode curly quotes (""') in the original text dataset. The GPT-1 paper⁠ mentions using the ftfy Python library to clean up mojibake & Unicode in their crawl data, and ftfy converts Unicode quotes to the ASCII straight quotes, so presumably GPT-2 does as well and it (or its BPE encoding) is confused by their presence, causing the mojibake output. It was late in training, but we updated the PG+PF dataset to replace the quotes (ftfy.fix_text('foo') etc).
One issue worth noting was the problem of regularly restarting the swarm due to preemption & TPUs expiring, which caused large loss spikes on startup that would waste hours of training as it recovered; as a compromise between simple SGD and full Adam, we were using Adafactor⁠ (as most Transformer projects do, like Connor or Gokaslan’s GPT-2 replications), and we speculate that the loss spike is related to losing optimizer state and bad initial variance estimates. Simple SGD avoided the loss spike, but at the cost of making no discernible progress regardless of LR; Adafactor made slow progress, but wasted a substantial fraction of available training time; we tried to avoid Adam because the memory overhead of tracking momentum for all variables (as opposed to Adafactor’s simplified approximation of momentum) would reduce minibatch size, but when we tried Adam on the full swarm, despite the initial loss spike, it made much more rapid progress than Adafactor did.
I suspect that the issue here is that though simple SGD & Adafactor worked fine when running on a single GPU, in the scaled-up asynchronous swarm setting, they have especially poor gradient estimates and make slow progress; the loss spike comes from the optimizer state being reset on startup, causing early gradients to be poorly estimated & destabilizing training, requiring thousands of iterations to gradually recover. Adam then improves over Adafactor by estimating true variance/momentum, overcoming gradient noise to make faster progress. If so, the spike issue could be fixed several ways:
  1. Don’t Reset The Optimizer: the simplest way to fix the spike caused by resetting momentum estimates is to not reset them; save the optimizer state along with the model, and restore on startup. This is somewhat unusual for TensorFlow projects (I see more PyTorch implementations serializing optimizer state) but shouldn’t be hard, and only comes at the cost of using more diskspace for checkpoints & more bandwidth at startup, which would be more than worthwhile to save hours of recovery time.
  1. Learning Rate Warmup: if the initial updates are highly destructive because they use bad momentum updates and it takes a substantial number of iterations to re-estimate the correct momentum updates, then don’t update much initially; use small LRs to avoid destabilizing the swarm while still re-learning the momentum. After some iterations, the updates should be safe to use again and the LR can be increased to the normal LR.
  1. Gradient Accumulation: another way to reduce the damage of initial updates is to greatly improve their accuracy, and keep the swarm more in sync, by doing fewer but better updates; instead of each node doing local updates immediately after each tiny local minibatch, the nodes store gradients and average across many minibatches before doing an actual update, thereby faking having large updates. This will also improve the momentum estimates, as the momentum is estimated across many minibatches before it affects any update.
Partway through, having reached a loss of ~2.6 (down ~0.5 from the Colab model), we experimented with training our model on a P100 GPU, halving the context window to make it fit, to informally compare its training speed with the swarm. The P100 made little training progress, but it did generate some fun poetry samples (we had disabled the training sample generation for the swarm because generating samples is so slow).
The samples strike me as good, perhaps even better than GPT-2-117M, despite the loss being much worse (2.6 rather than 1.6). Why might that be?
I hypothesize it reflects a weakness of the likelihood loss in terms of perceptual quality: humans are more sensitive to long-range correlations and text degenerating into gibberish than we are to local details like exact use of particles or to slightly better modeling of spelling (which is why stylometrics⁠ works). The original OA GPT-2-1.5b achieves much better modeling of long-range correlations and producing coherent text than the GPT-2-117M did, of course. What happens when they are both trained on a poetry dataset? It is the tale of the tortoise & the hare, or the bias-variance tradeoff⁠: the GPT-2-117M is weak, bad at long-range modeling because of its small parameter count & shallow layers, but the benefit is that it can learn quickly about local details like spelling, and, achieving good prediction there, converge to that 1.6 loss; GPT-2-1.5b starts off good at long-range modeling and good at short-range modeling, and must tradeoff learning both from its limited training, thereby achieving mediocre performance on local correlations and thus mediocre loss, even though humans reading it are impressed by the thematic consistency and relative lack of ‘gibberish’ (locally but not globally consistent text).
An additional issue here is that the GPT-2 models are not fully trained: as ⁠the GPT-2 paper notes⁠, “All models still underfit WebText and held-out perplexity has as of yet improved given more training time.” (The difficult of training such powerful LMs to convergence was also noted by the MegatronLM researchers⁠, whose MegatronLM-8.3b model was still learning rapidly when they ended the training—despite use of NVIDIA’s DGX SuperPOD with 512 GPUs.) So some of the finetuning here may also be finishing the GPT-2 training.
In the Golden Age, when the people of the Yellow Valley were instructed by the sages of antiquity:
Romance?
A nice descriptive piece:
An elegy:
Love lost:
An attempt at nonsense verse, apparently:
The wreck of a ship:
GPT-2-1.5b can apparently do meta-fiction and break the fourth wall‽
Another shipwreck:
A surprisingly coherent piece on a trapped upper-class wife:
The curse of immortality:
Perhaps the most striking of them all is this existential horror piece:
The expanded TPU swarm & Adam LR tuning allowed rapid training, and we reached 1.6 overnight, matching our previous best on the combined PG+PF poetry dataset.
Criticism of England:
The world of the dead:
The Demiurge:
Jealousy:
War. War never changes:
The secret garden:
I don’t really get this one but the repetition and inversions make it interesting to read:
Love lost:
Urban vs rural life:
Social media satire:
Pantheism:
Art:
An elegy?
One more I noticed while ⁠generating from a ~1-loss model on 2019-12-16⁠, although I did not read through for selections (see also ⁠~0.7-loss model samples, 2019-12-18⁠):
Subjectively, the output shows a lot of poetry knowledge, much better than the char-RNN samples. There’s (some) rhyming, themes are continued for shockingly long passages compared to char-RNN, and there are many passages I feel could inspire a poet or even be cleaned up a little to be passable poems on their own. Adding the metadata did help—GPT-2-poetry is worse than GPT-2-poetry-prefix. Some of the ones I liked most are (first lines) ‘We never say “Thank you”’, ‘Thy soul, thy very soul is burning!’, ‘“It is morn!” said the clover-bush’, ‘And they have seen the last light fail’, ‘There comes a murmur low and sweet’, and probably the best is ‘The sun is gone, and the night is late’.
Is GPT-2-poetry-prefix better than GPT-2-117M at poetry completions (since GPT-2-117M will probably hardly ever generate poetry without a prompt)? Probably, with exceptions. “Howl” is far worse, but that is for good reason related to the oldness of the PG corpus; if anyone could assemble an equally large corpus of more recent poetry, I’d expect GPT-2-117M finetuning to produce better completions. The Pope samples from GPT-2-poetry-prefix are clearly better (before diverging into prose). I argue that the Shelley samples are somewhat better. And the 8 famous line completions are overall of much higher poetic quality (several of the GPT-2-117M completions are just prose, unsurprisingly).
So, if one is looking for poetry completions in an old-fashioned vein, it delivers, but at the cost of flexibility like more prose-like (and hence contemporary) poems. This is an expected and fixable problem, and overall, I consider GPT-2-poetry-prefix to be successful as a poem generator & better than my previous char-RNNs.
Nor is this near the limit for Transformer-based poetry generation, as there are many possible improvements which could be made, all of which I’d expect to deliver substantial gains:
  • Make It Bigger:
    • bigger NN models: our initial results used the publicly-released GPT-2-117M, which delivers inferior results on all tasks compared to the unreleased GPT-2-1.5b: the samples generated by OpenAI & associates from GPT-2-1.5b are much better than GPT-2-117M samples, indicating that simply scaling up continues to deliver gains. Our ⁠GPT-2-1.5b⁠ poems turned out substantially better.
      • Nor did the various GPT-2 model sizes appear to reach any natural limit with GPT-2-1.5b, indicating that the Transformer NNs can be increased much further before hitting zero marginal gains. (This is consistent with other large-scale NN research, particularly on CNNs⁠ where even billions of images can be usefully trained upon.) OpenAI’s ⁠Greg Brockman has said⁠ (February 2019) that OpenAI intends to keep scaling GPT-2-1.5b up with aspirations of training 10–1000‘GPT-2-huge’ and a 1000× bigger still ‘GPT-2-enormous’ is possible, the quality leap from GPT-2-117M poetry to a hypothetical ‘GPT-2-enormous’ would be staggering.
        These projections for GPT-3 have since been borne out—GPT-3 (published 2020-05-28) has 175b parameters (166× more), and ⁠GPT-3’s untrained random poems⁠ are as good or better (!) than our GPT-2-1.5b poems.
    • better NN models (which will probably need to be bigger): the most painful limit is the small context window, which has a number of possible solutions⁠ like recurrency, memory, efficient attention variants, or various approximations. other options include more attention heads or more layers or external memory functions or on-the-fly adaptation; there are many possibilities here. (The prefix can be seen as an extremely crude kind of recurrency or memory, and helped a lot; how much more so a real memory?)
    • more & better data: quantity-wise, the PG corpus is barely a tenth of a gigabyte and exhibits many enormous omissions—all of modern poetry, for example, not to mention most foreign poetry, or non-English poetry as a whole (why not a multi-lingual GPT-2 if sufficiently large? neural machine translation approaches improve the more languages they have access to, why not regular language generation?). There are many places additional poetry could be obtained from, such as WikiSource, Poetry Foundation, Libgen, or the Internet in general (perhaps write a poetry-detector Transformer to search through a dump like Common Crawl⁠ for poetry?). Quality-wise, the PG corpus is good but still has a number of flaws: a lot of prose, just enough non-English poetry to screw things up (especially Latin), mostly pre-1923 poetry, & minimal metadata (ideally, poems would be individual units rather than book-length streams, and metadata like author would be available to use in prefixes).
      • GPT-3 expands the dataset greatly to more of Common Crawl plus 57 billion words from 2 deliberately-vaguely-described Internet book corpuses, but playing with it, I feel that GPT-3 is still weak on areas like science—for example, I observe GPT-3 to easily write machine learning paper abstracts, but it does not do quite so well when I try to extend it to the rest of papers, and it doesn’t spontaneously quote from within papers the way it quotes abstracts. Does this reflect the fact that many papers exist only as PDFs, and only a relatively small fraction of all scientific papers have clean readable HTML versions (while they usually all have readable HTML abstracts)? If so that may weaken the GPT reasoning & common sense abilities considerable; after all, while it does not usually come up in regular writing that giraffes have two eyes instead of three eyes, probing definitions & studying exceptions & manipulating causal arrows in unusual ways are all the bread & butter of scientific writing, and could implicitly teach that better than regular writing.
  • Generate Smarter
    • using a better sampling strategy than top-k, like “nucleus sampling”⁠ (but curiously, not beam search⁠—beam search gives substantial improvements on what the nucleus sampling authors call “closed” text generation tasks like translation, but while beams search helps char-RNN a little⁠, it damages results badly the wider the beam, and gives particularly bad results on GPT-2; ⁠Kyle Kastner says⁠ that beam search can work in contexts with heavy constraints, like being constrained to generate rhyming lines or explicit repetition penalties; Douglas Summers-Stay⁠ overcomes the rhyming limits by brute force: generating random completions until target lines rhyme according to a rhyming dictionary library)
      • Nucleus sampling has been implemented in nshepperd’s Tensorflow & Hugging Face’s PyTorch GPT-2 sampling code.
    • use tree search methods: any deep, thorough, search inevitably becomes a tree; tree searches are useful for enabling kinds of ‘backtracking’ and ‘revision’ or ‘changing its mind’ about multiple possible variants of a poem, as opposed to the usual sampling approaches which tend to commit to each word and force all-or-nothing choices. (My proposal for backprop reward optimization⁠ would have similar advantages, as each iteration step allows ‘thinking’ about how to improve a given input, approximating a search implicitly—even if not explicitly like a MCTS⁠ or MuZero⁠like approach.) The challenge here is to figure out a tree search which avoids the repetition trap.
  • Train Better, by fixing the loss (eg. unlikelihood training⁠⁠⁠18⁠), or switching to the RL setting to directly maximize generation quality:
    • richer losses: the standard GPT unidirectional prediction loss is not the only possible (differentiable⁠) loss; it is not even, strictly speaking, the best—models like BERT⁠/BART⁠ using more sophisticated losses like bidirectional losses, which force the model to predict a word missing from anywhere in the string (as opposed to only missing from the end), typically outperform GPT-2 on language tasks. A model like T5⁠ uses a denoising objective where a 15%-long chunk is replaced by a missing token & T5 must predict all the missing text based on context; these sorts of objective losses allow learning much more from a given dataset. (Indeed, such models typically outperform GPT-2 on everything but language generation. Oddly, they typically do quite badly at that, which is a major reason everyone uses GPT-2 for generating new texts, and BERT etc for everything else like generating embeddings or classification.)
    • adding global end-to-end losses, which enable training to optimize non-differentiable properties rather than easy (but partially irrelevant ones like predictive losses such as cross-entropy in prediction of the next word). For example, rules defining acceptable meter or rhyme use or penalizing total repetition—these cannot be done via the normal training because no individual discrete word is responsible and parameters cannot be smoothly adjusted to decrease/increase a global property like ‘rhymes’ which is the result of all words considered together as a whole. (This sort of RL loss has been employed in other natural language tasks like machine translation, where metrics like predictive loss do not map onto the desired goal of semantically-correct translation, and word-by-word generation of translations yields similar issues as here, but there are metrics like BLEU⁠ or ROUGE⁠ or grammar checkers which provide a crude measure of global quality. RL approaches have many virtues⁠.)
    • instead of training a NN to predict individual next-characters as accurately as possible or imitate a text corpus as well as possible, we really just want them to predict good next-characters to write text as well as possible—which is not the same thing at all, any more than accurately predicting a human Go player’s next move on average is the same thing as playing Go superhumanly well.
      • This encourages more global coherency, more thematic progressions, use of rare words when appropriate, surprising subversions or twists which work well when tried but don’t appear in the original corpus, learning esthetics, and so on. If it works and the new GPT-2-poetry is able to successfully produce new poems which consistently get the top score from the critic and no further improvement is happening, then you simply read a bunch of its new poems, pick which one in each pair you like, retrain the critic on the expanded dataset to detect the remaining flaws in the ones you disliked, and then keep training GPT-2-poetry to avoid generating the ones you disliked & generate more poems like the ones you liked. Repeat with many cycles, and it should generate excellent poems while avoiding all the flaws of crude likelihood training and even cruder top-k sampling which hobble GPT-2-poetry right now. Even better, you could create a website to crowdsource the rankings to keep it training 24/7 and improving indefinitely.
    • using “expert iteration” architectures like AlphaZero⁠ to do much more sophisticated search over possible poems, creating an iterative bootstrap
    • adding creativity losses along the lines of “CAN: Creative Adversarial Networks, Generating ‘Art’ by Learning About Styles and Deviating from Style Norms”⁠, Elgammal2017, where updating GANs encourage diversity
      • one could attempt to invent new styles of poetry by taking inspiration from evolutionary methods, such as the “Population-Based Training” variant employed in ⁠DeepMind’s AlphaStar League⁠ which created diversity by deliberately scrambling the ‘rules’ for each lineage of agents. The “AlphaStar League” used a population of multiple NNs, each forced to specialize in using a particular unit or rewarded for achieving particular goals like defeating a specific NN (rather than winning in general). The AlphaStar League was credited for forcing the overall AlphaStar population to explore strategies reliant on particular kinds of units and figuring out counter-strategies to successful ones. Something similar could be done with poetry rules: train many different agents, each given a specific rhyme scene or meter or vocabulary for their reward function, and in preference-learning approaches, the best poems can be provided to human critics for rating & improving the NN critic. Potentially exciting new combos could emerge as producing the best poems as rated by the humans.
Given that GPT-2-117M is far from the state-of-the-art as of February 2019, and hardware & generative NN research is advancing rapidly, it will be exciting to see what sort of poetry can be generated given another 4 years!
  • Discussion:
  • poem-generator⁠ (Generates rhyming poetry using Huggingface GPT-2 using rejection sampling—throws away possible completions which don’t rhyme)
  • lm-scorer⁠ (“This package provides a simple programming interface to score sentences using different ML language models.”)
  • “Excavate”, Mike Lynch (search over a corpus to extract a RNN-generated ‘hidden text’)
Aaron Gokaslan scraped the large fanfiction website Archive of Our Own⁠ (Ao3) and created a text dump⁠ (2.7GB archive, 12GB raw; 190,931 stories; 2.06b words; no metadata).⁠⁠19⁠
We trained a GPT-2-1.5b on it (checkpoint⁠, 11GB), under the theory that it might be useful as a basis for text-game-like applications such as AI Dungeon 2, under the idea that since AI Dungeon 2 is essentially collaborative story-telling, starting with a story-based model ought to give better results.
⁠AstraliteHeart⁠ has (w/Shawn Presser) trained a finetuned GPT-2-1.5b on a large scrape of fiction as part of work on a forthcoming MLP video voice-chat dialogue tool, PurpleSmart.⁠⁠20⁠ The fiction sources were:
(Total data size: 200GB compressed pre-filtering, currently unclear how much it was actually trained on after quality filtering; because of the wide variety of fiction, we dub it ‘uberset’.)
The accompanying Tacotron2-Torchmoji & WaveGlow models can likewise be downloaded⁠.
Twitter user me_irl⁠ provided a 50MB scrape of video game walkthroughs, which he’d used previously with GPT-2-345M and requested we do finetuning on that as well: ⁠video game walkthrough text samples⁠ (Newsweek article⁠). me_irl has suggested that they could be used as hypothetical game designs for competitions or art purposes.
In December 2019⁠, Shawn Presser trained a GPT-2-117M model for a few million steps on the /r/DoTA2 subreddit, as part of the Subreddit Simulator training project. (The final GPT-2-1.5b & dataset for Subreddit Simulator has not been released at the project’s request.)
While the final model checkpoint appears to have been lost (oops), step #562,971 from 2019-12-18 has been uploaded:
  1. A Transformer is a considerably different architecture than an RNN, and is not that easy to explain, as it uses multiple convolutions to implement “attention”, allowing flexible internal control flow, over a large but finite input window, without any recurrency or hidden state or LSTM⁠ units necessary. For increasingly-technical explanations, see:
  1. 774M requires changes to nshepperd’s checkpointing, specifically, removing the layer == 10 restriction in model.py, and letting the checkpointing code checkpoint as much as possible, which enables training minibatches n≤10 on my 2×1080tis. Diff:
    1. It would require either high-end GPUs with ≥16GB VRAM, or TPU instances (which were used to train it). GPT-2-1.5b can’t be trained on my 1080tis with either the nshepperd codebase or Shawn Presser’s fork, although Presser has a Google Colab notebook using TPUs⁠ which can train it.
    1. Other examples of finetuning are ⁠Facebook Messenger logs, nshepperd’s unpublished Linux kernel C source code & IRC-log training⁠⁠4⁠, and story prompts. And, while it doesn’t use GPT-2-117M, too good to not mention is ⁠“Stack Roboflow: This Question Does Not Exist”⁠.
    1. GPT-2 completions of 26 prompts: “Ozymandias”/“One Art”/“The Road Not Taken”/“Where the Sidewalk Ends”/“Because I could not stop for Death”/“Inferno, Canto I”/“In Flanders Field”/“O Captain! My Captain!”/“Howl”/“The Tyger”/“Outsight”/“Zuang Zhou Dreams of Being a Butterfly”/“Sonnet”/“Oh, the Places You’ll Go!”/“The Hollow Men”/“The Summer Day”/“A Just-Finishing Candle”/“A Psalm of Life”/“Still I Rise!”/“The Second Coming”/“Do not go gentle into that good night”/“Kubla Khan”/“Edge”/“The Raven”/“There Will Come Soft Rains”/“The Lorax”.
    1. For example, 2 people finetuned GPT-2-117M on an IRC channel’s logs, getting losses of 1.95 & 2.3; why was the latter’s loss 18% worse compared to the former when they were using the same IRC channel, GPT-2-117M pretrained model, training codebase, & had both apparently converged? Because. while the IRC channel was the same, they used different IRC clients which had different IRC log formatting conventions—the former’s logs had the full timestamp prefixed to each line, and the latter didn’t. Said timestamps made up ~20 characters of ~110 character lines, or, ~18% of each line! So the models were performing identically on the content that mattered, and the much lower loss was simply because of near-perfect prediction of the highly-repetitive & predictable timestamps on every line. (Indeed, given the limited window of GPT-2-117M, arguably the model with the worse loss would be better in terms of generating fun coherent samples.)
    1. I have 2 GPUs but nshepperd’s code does not (yet) support multi-GPU training easily. Some support using Horovod for multi-GPU has been added but I cannot vouch for it.
    1. I discovered this while being puzzled why -batchsize 32 did not lead to instant out-of-memory errors for training; similarly, if you make the mistake of sampling with the option -top 40, what you are actually doing is sampling with the default -top_k 0. Oops.
    1. It is possible that the additional training was helping, because the remaining tiny changes in the loss might translate to large perceived quality improvements—while the loss didn’t change, the samples from later on did strike me as better. This was something I thought I noticed with char-RNN as well, that the loss became a bad guide to quality when the NN had mostly converged. On the other hand, with larger GPT-2s, like GPT-2-1.5b, the relationship between loss and perceived quality seems even more opaque, with quality sometimes worsening even as the training loss decreases rapidly.
    1. One might worry that by taking up space in the model’s limited context ‘window’ of inputs, because the Transformer has no hidden state or ‘memory’, such inline metadata would be a bad thing as it will push real words out of the context window, thereby degrading quality and making it even more incoherent & rambling.
      1. But on the other hand, if it does learn to associate specific IDs with genres/topics, then repetition of the inline metadata serves as a ‘mnemonic’ for global information which is available to all subsequent iterations of the model, serving as a crude memory itself.
        For example, if Homeric pastiche has ID #16452, then as long as the final iteration of the model overlaps for just the ID with the first iteration of model during sampling and both see “16452”, all models will be able to consistently agree on generating Homeric pastiche rather than some other pastiche because they all see the same ID somewhere in their context window & that guides their generation.
    1. ⁠starspawn0⁠ has collated some of the results:
      1. We also introduce (a) a variation on architecture and initialization to train deeper networks, (b) the recomputation of attention matrices to save memory, and (c) fast attention kernels for training. We call networks with these changes “Sparse Transformers”, and show they can model sequences tens of thousands of timesteps long using hundreds of layers. We use the same architecture to model images, audio, and text from raw bytes, setting a new state-of-the-art for density modeling of ⁠enwik8, CIFAR-10⁠, and ImageNet⁠64. We generate unconditional samples that demonstrate global coherence and great diversity, and show it is possible in principle to use self-attention to model sequences of length one million or more. Gradient checkpointing has been shown to be effective in reducing the memory requirements of training deep neural networks (Chen2016⁠), (⁠Gruslys2016⁠). It is worth noting, however, that this technique is particularly effective for self-attention layers when long sequences are processed, as memory usage is high for these layers relative to the cost of computing them.Using recomputation alone, we are able to train dense attention networks with hundreds of layers on sequence lengths of 16,384, which would be infeasible on modern hardware otherwise. In our experiments, we recompute the attention and feed-forward blocks during the backwards pass. …For each sequence length, we attempted to train the largest model which could entirely fit into 16GB V100⁠ accelerators without model parallelism. Overall, we found that increasing the sequence length by a factor of 4 requires a reduction in model capacity of approximately 4 × √4 = 8. Thus we found we could use factorized self-attention on sequences over 1 million timesteps long, albeit with extremely few parameters (3 million).
      1. The code turns out to multiply by a large number as a way of setting a default ‘highly unlikely’ value for each possible BPE, but in FP16, it can’t be represented and overflows, and so the output amusingly just becomes the BPE 0, which is the character ‘!’, so it kept printing out ‘!!!’. Indeed.
      1. The Colab environment has special support for mounting Google Drive, so a magical incantation like this will mount a Google Drive folder as a normal (slow) directory:
        1. Cloud provider bandwidth like Amazon AWS or GCP are notoriously high for “egress” traffic leaving the cloud provider; like Hotel California, they want you to check in but never leave, and charge egressicous fees per gigabyte. This is why I try to use my Hetzner dedicated server for all hosting, which will let people download terabytes without bankrupting me.
        1. We noticed that, in fact, our preemptible TPUs would always preempt precisely at midnight. I speculated that, as discussed in ⁠the Google SRE handbook⁠ where they cover how the Chubby service is deliberately taken down at random to live down to its uptime promises, preemptible TPUs were being deliberately taken down to stop users from treating them like on-demand TPUs, and this simply wasn’t documented. Discussing our problems with TRC, this apparently was correct.
        1. Presser is convinced that TPU power is greatly overrated and most TPU projects have made poor use of the available power, as they get far less speedups than one would expect over GPUs. Oddly, there doesn’t seem to have been much work on using multiple TPUs outside of a TPU pod configuration, and TRC did not seem to know of an equivalent to the ‘swarm’.
        1. See2019⁠ also demonstrates problems with likelihood decoding strategies in GPT-2 for story generation.
        1. Alternative Ao3 fanfiction dumps, among others, are available on ⁠the Internet Archive⁠.
        Loading...