Tuesday, June 13, 2017

Feminine and Masculine Rhymes in Debosnys' Poem

I've been reading a lot of Beaudelaire's poetry over the last couple of days, in order to better understand the structure of French poetry. (I picked Beaudelaire because I think there is a fair chance that Debosnys might have been influenced by him.)

One of the fairly consistent features of French rhyme in Beaudelaire, as well as in Debosnys' plaintext poem, is the alternation of masculine and feminine rhymes. A feminine rhyme is a rhyme ending in the silent e. In Debosnys' plaintext poem, the feminine rhymes are: grâce / passe, blessure / dure, brave / nuage (!).

I think there is a high likelihood, if the cipher poem is in French, that it alternates masculine and feminine rhymes. Weak supporting evidence for this is the fact that the only repeating rhyme occurs on the odd-numbered couplets 1 and 9. Since that rhyme would either be masculine or feminine, it should occur only on either odd or even lines.

There is nothing obvious that indicates which lines are masculine and which are feminine. There are two odd rhymes that use symbols that are variations on ♀, and one even rhyme that uses a symbol that is a variation on♂, but that is far from conclusive.

Sunday, June 11, 2017

Minor update on the Debosnys poem

I've done some reading on French poetry and some work on transcribing the Debosnys cipher texts, and found three things that are interesting.

1. Identical rhyme

Conventional French rhyme apparently often includes the onset of the syllable in the rhyme, so cœur rhymes with vainqueur, and sonores rhymes with Maures.

If the rhyme includes the onset of the syllable, then the rhyming symbols of the text could quite easily correspond to full syllables, rather than needing to be broken at the nucleus of the syllable.

Debosnys did not use this type of rhyme in his plaintext poem, but that is no guarantee that the cipher poem does not use it. It is also possible that the text of the cipher poem was not authored by him. For all we know, he may have enciphered a poem by some other author.

2. Irregular Meter

I initially thought Debosnys' plaintext French poem was in iambic pentameter, but that was only on the basis of the first four lines. When I counted syllables for the rest of the lines, I found a fair amount of variation. Some of the other lines have the 12 syllables of an alexandrine, but there is no strict rule, and the longest line has 14-15 syllables.

If the cipher poem was authored by him, then the variation in the number of symbols per line could correspond to a variation in the number of syllables.

3. At least two languages

I have transcribed the cipher poem and one block of text (appearing above a plaintext poem in French), and I think they use the same cipher to write two different languages. The symbol frequency and inventory are quite different. Apparently Debosnys was a bit of a polyglot, so I don't know which two languages these are.

Friday, June 9, 2017

First hypothesis for the Debosnys cipher

In my last post I noted that the last symbol on each line of the Debosnys poem seems to represent the whole rhyme and nothing but the rhyme. To me, this says three things:

  1. The cipher is phonetic
  2. Some symbols represent rhymes.
  3. The onsets of syllables are either left out or encoded in some other way

The first thing I thought of when I saw this was a type of pseudo-language used in the English version of Hergé's book Tintin and the Picaros. (I wish I had a copy of the French original to compare, but alas, I don't).

The Picaros use a language with phrases like the following:

Goh'blimeh! wa'samma ta, li li li va? Lem eshohya!
Sum in'ksup wivit!

The words are accented English, spelled phonetically, and broken in unnatural places. The process of obfuscation could be described in three steps:

Plain Text: "Something's up with it"
Accented Phonetic: sumink's up wiv it
Broken: sum in' ks up wiv it
Re-merged: sum in'ksup wivit

One of the Debosnys pages shows a poem in French that is roughly in iambic pentameter, with an ABAB rhyme scheme. (This strikes me as the product of an English speaker who learned French as a second language, since iambic pentameter isn't common in modern French verse...though of course it was common on Old French.)

Suppose we take the first two lines of Debosnys' plaintext French poem, represent them phonetically, and break them up into groups representing the rhymes of each syllable combined with the onset of the next:

oh! mes amis je vous supplie en grâce
de bien vouloir un instant m'écouté

Step 1: Represent it phonetically. I'll do that using an 1880 phonetic dictionary:

ō mé-z-ămī jĕ vǒû süplī ĕñ grâs
dě bĭĕñ vǒûlwâr üǹ ăñstăñ m-ékŏûté

Step 2: Break at the onset of rhyme

|ō m|é-z-|ăm|ī j|ĕ v|ǒû s|üpl|ī |ĕñ gr|âs
d|ě b|ĭĕñ v|ǒûlw|âr |üǹ |ăñst|ăñ m-|ék|ŏût|é

Step 3: Merge

ōm éz ăm īj ĕv ǒûs üpl ī ĕñgr âs
d ěb ĭĕñv ǒûlw âr üǹ ăñst ăñm ék ŏût é

Suppose Debosnys encoded each group of this text as a separate symbol. For a line of iambic pentameter, the product of this type of process would normally have 10-11 symbols. The lines of the cipher poem have, on average, 15 symbols per line. If the cipher poem is in iambic pentameter, then presumably complex or uncommon groups (like ĕñgrǒûlw) could be represented as two symbols, for example:

ōm éz ăm īj ĕv ǒûs üp l ī ĕñ g r âs
d ěb ĭĕñ v ǒûl w âr üǹ ăñ s t ăñ m ék ŏût é

Thursday, June 8, 2017

A shiny thing

I was looking for substitution ciphers to test my new tools on, and I discovered one I hadn't heard about, thanks to Nick Pelling's Cipher Foundation site: The Debosnys Cipher.

Among the Debosnys cipher pages (the images of which I got from Pelling's site) there are the following two, which are clearly a poem.

It is appears from looking at the text that these are rhyming couplets with a scheme of AA BB. Assuming this is a conventional European end-rhyme, the rhyme should consist minimally of the nucleus and coda of the last syllable of the word. Since the visual rhyme appears to be limited to the last grapheme of each line, that grapheme seems to represents the whole rhyme and only the rhyme.

Suppose each symbol represents either an onset or a rhyme. In that case, we would expect the rhyming symbols at the ends of the lines to be slightly more frequent in the second position on each line than random chance would allow, since the first symbol would presumably represent an onset, and the second a rhyme. Indeed, out of 20 lines, four have end-rhyme symbols in the second position (20%), compared to one in first position (5%) and one in second-to-last position (5%).




Wednesday, June 7, 2017

First pass at mapping one network onto another

I have written some code to map networks onto each other, to try to find the best match between the topologies of the two networks.

It could run for a very long time, depending on the size of the network. Using the top 20 words from chapters 81 and 87 of Moby Dick, and running it for 8.5 minutes, I get a fairly good initial set of matches:


The groups in this graph represent the matches established between words in the two chapters. For six words (the, for, and, of, was, with, this) I got exact matches. The remaining groups were incorrectly matched after running for 8.5 minutes.

Even data of this quality could be useful. Suppose you have a text that must be in a known language, but you don't know what language it is. You could try all of the likely languages and generate a set of possible matches between words in your text and words in the candidate languages.

Tuesday, June 6, 2017

How to match word cipher text against reference text

In my last post, I mentioned that I want to develop some strategies to attack word ciphers. The reason for this is that I think it would be generally useful for decipherments like the Rohonc text, where the semantic domain is known.

One way to do this would be to create a network showing the relationships between words within a cipher text, and try to find the best match between that network and a similar network for a known text.

I am trying these ideas out with chapters 81 and 87 of Melville's Moby Dick. You can see the basic idea if you look at the closest relationships between the top 20 words of each chapter:



Closest relationships between top 20 words in Chapter 81


 

Closest relationships between top 20 words in Chapter 87

You can see that in both chapters there is a little island of nouns (whale, whales, it, he), and another little island of determiners (a, the, his). Most prepositions are connected to other prepositions, and the pronouns "that" and "this" are connected to each other. In the underlying data there is much more information available about the closeness of the relationships, but that is not shown in these graphs.

The main problem I will run into is the problem of processing time. Luckily, in the quiet years since I was last writing about these things, I have learned to use cloud computing. It will just be a big job and it needs to be planned out carefully.

Tuesday, May 16, 2017

Quiet a long time

I've been quiet a long time, because I've been very, very busy with work, kids and other things.

I find my best free time these days is when I am sitting in a dark room getting kids to fall asleep, but it's hard to work on the kinds of projects I used to do under those conditions. Instead I've been revisiting languages that I've studied in the past, and learning about some new ones.

But I've also been trying to figure out how I can fit more interesting work into my free fragments of time. I've built a kind of operating system that runs in a browser, that I can bring up on my phone or any laptop I happen to be using, and pick up a piece of work and poke away at it. The state of the session is saved to a server, and reloaded in the next browser I use to access it. I call it Joss.

One of the first applications I wrote to run on Joss is a JavaScript REPL, but it turns out to be fairly difficult to write JavaScript code on my phone. Autocorrect wants to fix everything up, and special characters require several taps to access. My next project is going to be a programming language that is suited to the constraints of the phone. I've been thinking about the properties of an English-like programming-language for years, and this is the perfect opportunity to try it out.

Ultimately, I want to get to a place where I can work on some other ideas that I have been gnawing on in my spare time, too. For example, I want to develop some strategies to attack word-substitution ciphers. I plan to take a text and divide it into two parts, and encipher one half with a randomly generated word substitution cipher, then use the other half as data to inform the decipherment. I want to use two halves of a single text in order to maintain a common semantic domain and idiolect.

Another idea I want to work on is something I call "paleo-poetics". This is inspired by work I did with Manchu to reconstruct syllable structure and prosody by looking at poetry. I think it would be interesting to apply the same ideas to Egyptian, Akkadian and Sumerian poetry, to see if I can discover anything interesting about those languages and their poetic traditions by uncovering the skeletons of their poetic forms.