Probing the Statistical Properties of Unknown Texts: Application to the Voynich Manuscript

PLoS ONE 8 (7):e67310 (2013)
  Copy   BIBTEX

Abstract

While the use of statistical physics methods to analyze large corpora has been useful to unveil many patterns in texts, no comprehensive investigation has been performed investigating the properties of statistical measurements across different languages and texts. In this study we propose a framework that aims at determining if a text is compatible with a natural language and which languages are closest to it, without any knowledge of the meaning of the words. The approach is based on three types of statistical measurements, i.e. obtained from first-order statistics of word properties in a text, from the topology of complex networks representing text, and from intermittency concepts where text is treated as a time series. Comparative experiments were performed with the New Testament in 15 different languages and with distinct books in English and Portuguese in order to quantify the dependency of the different measurements on the language and on the story being told in the book. The metrics found to be informative in distinguishing real texts from their shuffled versions include assortativity, degree and selectivity of words. As an illustration, we analyze an undeciphered medieval manuscript known as the Voynich Manuscript. We show that it is mostly compatible with natural languages and incompatible with random texts. We also obtain candidates for key-words of the Voynich Manuscript which could be helpful in the effort of deciphering it. Because we were able to identify statistical measurements that are more dependent on the syntax than on the semantics, the framework may also serve for text analysis in language-dependent applications.

Links

PhilArchive



    Upload a copy of this work     Papers currently archived: 91,881

External links

Setup an account with your affiliations in order to access resources via your University's proxy server

Through your library

Similar books and articles

The work of E. T. Jaynes on probability, statistics and statistical physics.D. A. Lavis & P. J. Milligan - 1985 - British Journal for the Philosophy of Science 36 (2):193-210.
Probabilistics: A lost science.L. S. Mayants - 1982 - Foundations of Physics 12 (8):797-811.
Insights in How Computer Science can be a Science.Robert W. P. Luk - 2020 - Science and Philosophy 8 (2):17-46.
Essay review: Probability in classical statistical physics.Janneke van Lith - 2001 - Studies in History and Philosophy of Science Part B: Studies in History and Philosophy of Modern Physics 33:143–50.
Speed of computation and simulation.Subhash C. Kak - 1996 - Foundations of Physics 26 (10):1375-1386.
Entropy in operational statistics and quantum logic.Carl A. Hein - 1979 - Foundations of Physics 9 (9-10):751-786.

Analytics

Added to PP
2023-09-18

Downloads
2 (#1,804,618)

6 months
1 (#1,471,540)

Historical graph of downloads

Sorry, there are not enough data points to plot this chart.
How can I increase my downloads?