Movatterモバイル変換


[0]ホーム

URL:


Jump to content
WikipediaThe Free Encyclopedia
Search

Baudot code

From Wikipedia, the free encyclopedia
Pioneering five-bit character encodings

An early "piano" Baudot keyboard

TheBaudot code (French pronunciation:[bodo]) is an earlycharacter encoding fortelegraphy invented byÉmile Baudot in the 1870s.[1] It was the predecessor to the International Telegraph Alphabet No. 2 (ITA2), the most commonteleprinter code in use beforeASCII. Eachcharacter in the alphabet is represented by a series of fivebits, sent over acommunication channel such as a telegraphwire or aradio signal byasynchronous serial communication. Thesymbol rate measurement is known asbaud, and is derived from the same name.

History

[edit]

Baudot code (ITA1)

[edit]
Baudot code (ITA1)
An early version from Baudot's 1888 US patent, listing A through Z,t and ∗ (Erasure)
Alias(es)International Telegraph Alphabet 1
Current statusReplaced byITA2 (not mutually compatible).
Classification5-bitstateful[citation needed]basic Latin encoding
Preceded byMorse code
Succeeded byITA2

In the below table, Columns I, II, III, IV, and V show the code; the Let. and Fig. columns show the letters and numbers for the Continental and UK versions; and the sort keys present the table in the order: alphabetical, Gray and UK

Baudot code (Continental and UK versions).[2]
Europesort keysUKsort keys
VIVIIIIIICon­ti­nen­talGrayLet.Fig.VIVIIIIIIUK
---
A1A1
É&/1/
E2E2
IoI3/
O5O5
U4U4
Y3Y3
B8B8
C9C9
D0D0
FfF5/
G7G7
HhH¹
J6J6
FigureBlankFig.Bl.
ErasureErasure**
K(K(
L=L=
M)M)
NN£
P%P+
Q/Q/
RR
S;S7/
T!T²
V'V¹
W?W?
X,X9/
Z:Z:
t..
BlankLetterBl.Let.

Baudot developed his first multiplexed telegraph in 1872[3][4] and patented it in 1874.[4][5] In 1876, he changed from a six-bit code to a five-bit code,[4] as suggested byCarl Friedrich Gauss andWilhelm Weber in 1834,[3][6]with equal on and off intervals, which allowed for transmission of the Roman alphabet, and included punctuation and control signals. The code itself was not patented (only the machine) because French patent law does not allow concepts to be patented.[7]

Baudot's 5-bit code was adapted to be sent from a manual keyboard, and no teleprinter equipment was ever constructed that used it in its original form.[8] The code was entered on a keyboard which had just five piano-type keys and was operated using two fingers of the left hand and three fingers of the right hand. Once the keys had been pressed, they were locked down until mechanical contacts in a distributor unit passed over the sector connected to that particular keyboard, at which time the keyboard was unlocked ready for the next character to be entered, with an audible click (known as the "cadence signal") to warn the operator. Operators had to maintain a steady rhythm, and the usual speed of operation was 30 words per minute.[9]

The table "shows the allocation of the Baudot code which was employed in theBritish Post Office for continental and inland services. A number of characters in the continental code are replaced by fractionals in the inland code. Code elements 1, 2 and 3 are transmitted by keys 1, 2 and 3, and these are operated by the first three fingers of the right hand. Code elements 4 and 5 are transmitted by keys 4 and 5, and these are operated by the first two fingers of the left hand."[8][10][11]

Baudot's code became known as theInternational Telegraph Alphabet No. 1 (ITA1). It is no longer used.

Murray code

[edit]
Paper tape with holes representing the "Baudot–Murray Code". Note the fully punched columns of "Delete/Letters select" codes at end of the message (on the right) which were used to cut the band easily between distinct messages. The last symbols before the fully punched columns at the end are BRASIL CR LF CR FS (word Brasil, carriage return, line feed, carriage return, shift to figures)

In 1901, Baudot's code was modified byDonald Murray (1865–1945), prompted by his development of a typewriter-like keyboard. The Murray system employed an intermediate step: an operator used a keyboard perforator to punch a paper tape and then a transmitter to send the message from thepunched tape. At the receiving end of the line, a printing mechanism would print on a paper tape, and/or a reperforator would make a perforated copy of the message.[12]

Because there was no longer a connection between the operator's hand movement and the bits transmitted, there was no concern about arranging the code to minimize operator fatigue. Instead, Murray designed the code to minimize wear on the machinery by assigning the code combinations with the fewest punched holes to the mostfrequently used characters.[13][14] For example, the one-hole letters are E and T. The ten two-hole letters are AOINSHRDLZ, very similar to the "Etaoin shrdlu" order used inLinotype machines. Ten more letters, BCGFJMPUWY, have three holes each, and the four-hole letters are VXKQ.

The Murray code also introduced what became known as "format affectors" or "control characters" – theCR (Carriage Return) andLF (Line Feed) codes. A few of Baudot's codes moved to the positions where they have stayed ever since: the NULL or BLANK and the DEL code. NULL/BLANK was used as an idle code for when no messages were being sent, but the same code was used to encode the space separation between words. Sequences of DEL codes (fully punched columns) were used at start or end of messages or between them which made it easier to separate distinct messages. (BELL codes could be inserted in those sequences to signal to the remote operator that a new message was coming or that transmission of a message was terminated).

EarlyBritish Creed machines also used the Murray system.

Western Union

[edit]
Keyboard of ateleprinter using the Baudot code (US variant), with FIGS and LTRS shift keys

Murray's code was adopted byWestern Union which used it until the 1950s, with a few changes that consisted of omitting some characters and adding more control codes. An explicit SPC (space) character was introduced, in place of the BLANK/NULL, and a newBEL code rang a bell or otherwise produced an audible signal at the receiver. Additionally, the WRU or "Who aRe yoU?" code was introduced, which caused a receiving machine to send an identification stream back to the sender.

ITA2

[edit]
ITA2 Baudot–Murray code
British variant of ITA2
Alias(es)International Telegraph Alphabet 2
Classification5-bitstateful[citation needed]basic Latin encoding
Preceded byITA1
Succeeded byFIELDATA,
ITA 3 (van Duuren code),
ITA 5 (ISO 646,ASCII)
MTK-2
Language(s)Russian
Classification5-bitstateful[citation needed]Russian Cyrillic encoding
Preceded byRussian Morse code
Succeeded byKOI-7

In 1932, theCCITT introduced theInternational Telegraph Alphabet No. 2 (ITA2) code[15] as an international standard, which was based on the Western Union code with some minor changes. The US standardized on a version of ITA2 called theAmerican Teletypewriter code (US TTY) which was the basis for 5-bit teletypewriter codes until the debut of 7-bitASCII in 1963.[16]

Some code points (marked blue in the table) were reserved for national-specific usage.[17]

A four-row teletype keyboard with Roman and Cyrillic letters.
International telegraphy alphabet No. 2 (Baudot–Murray code)[18]
Impulse patterns
(1=mark, 0=space)
Letter shiftFigure shift
LSB on
right;
code elements:
543·21
LSB on
left;
code elements:
12·345
Count of punched marksITA2
standard
Russian
MTK-2
variant
Russian
MTK-2
variant
ITA2
standard
US TTY
variant
000·0000·0000NullShift to Cyrillic LettersNull
010·0000·0101Carriage return
000·1001·0001Line feed
001·0000·1001Space
101·1111·1014QЯ1
100·1111·0013WВ2
000·0110·0001EЕ3
010·1001·0102RР4
100·0000·0011TТ5
101·0110·1013YЫ6
001·1111·1003UУ7
001·1001·1002IИ8
110·0000·0112OО9
101·1001·1013PП0
000·1111·0002AА
001·0110·1002SС'Bell
010·0110·0102DДWRU?$
011·0110·1103FФЭ!
110·1001·0113GГШ&
101·0000·1012HХЩ£#
010·1111·0103JЙЮBell'
011·1111·1104KК(
100·1001·0012LЛ)
100·0110·0012ZЗ+"
111·0110·1114XЬ/
011·1001·1103CЦ:
111·1001·1114VЖ=;
110·0110·0113BБ?
011·0000·1102NН,
111·0000·1113MМ.
110·1111·0114Shift to Figures (FS)Reserved for
figures extension
111·1111·1115Reserved for
lettercase extension
Shift to Letters (LS)
/ Erasure / Delete

The code position assigned to Null was in fact used only for the idle state of teleprinters. During long periods of idle time, the impulse rate was not synchronized between both devices (which could even be powered off or not permanently interconnected on commuted phone lines). To start a message it was first necessary to calibrate the impulse rate, a sequence of regularly timed "mark" pulses (1), by a group of five pulses, which could also be detected by simple passive electronic devices to turn on the teleprinter. This sequence of pulses generated a series of Erasure/Delete characters while also initializing the state of the receiver to the Letters shift mode. However, the first pulse could be lost, so this power on procedure could then be terminated by a single Null immediately followed by an Erasure/Delete character. To preserve the synchronization between devices, the Null code could not be used arbitrarily in the middle of messages (this was an improvement to the initial Baudot system where spaces were not explicitly differentiated, so it was difficult to maintain the pulse counters for repeating spaces on teleprinters). But it was then possible to resynchronize devices at any time by sending a Null in the middle of a message (immediately followed by an Erasure/Delete/LS control if followed by a letter, or by a FS control if followed by a figure). Sending Null controls also did not cause the paper band to advance to the next row (as nothing was punched), so this saved precious lengths of punchable paper band. On the other hand, the Erasure/Delete/LS control code was always punched and always shifted to the (initial) letters mode. According to some sources, the Null code point was reserved for country-internal usage only.[17]

The Shift to Letters code (LS) is also usable as a way to cancel/delete text from a punched tape after it has been read, allowing the safe destruction of a message before discarding the punched band.[clarification needed] Functionally, it can also play the same filler role as the Delete code in ASCII (or other 7-bit and 8-bit encodings, including EBCDIC for punched cards). After codes in a fragment of text have been replaced by an arbitrary number of LS codes, what follows is still preserved and decodable. It can also be used as an initiator to make sure that the decoding of the first code will not give a digit or another symbol from the figures page (because the Null code can be arbitrarily inserted near the end or beginning of a punch band, and has to be ignored, whereas the Space code is significant in text).

The cells marked as reserved for extensions (which use the LS code again a second time—just after the first LS code—to shift from the figures page to the letters shift page) has been defined to shift into a new mode. In this new mode, the letters page contains only lowercase letters, but retains access to a third code page for uppercase letters, either by encoding for a single letter (by sending LS before that letter), or locking (with FS+LS) for an unlimited number of capital letters or digits before then unlocking (with a single LS) to return to lowercase mode.[19] The cell marked as "Reserved" is also usable (using the FS code from the figures shift page) to switch the page of figures (which normally contains digits andnational lowercase letters or symbols) to a fourth page (where national letters are uppercase and other symbols may be encoded).

ITA2 is still used intelecommunications devices for the deaf (TDD),Telex, and someamateur radio applications, such asradioteletype ("RTTY"). ITA2 is also used in Enhanced Broadcast Solution, an early 21st-century financial protocol specified byDeutsche Börse, to reduce the character encoding footprint.[20]

Nomenclature

[edit]

Nearly all 20th-century teleprinter equipment used Western Union's code, ITA2, or variants thereof. Radio amateurs casually call ITA2 and variants "Baudot" incorrectly,[21] and even theAmerican Radio Relay League's Amateur Radio Handbook does so, though in more recent editions the tables of codes correctly identifies it as ITA2.

Character set

[edit]

The values shown in each cell are theUnicode codepoints, given for comparison.

Original Baudot variants

[edit]

Original Baudot, domestic UK

[edit]
Original Baudot code, UK domestic variant (letter set, switched to with 0x10)[22]
0123456789ABCDEF
0xNULAE/YUIOFIGSJGHBCFD
1x SP -XZSTWVDELKMLRQNP
Original Baudot code, UK domestic variant (figure set, switched to with 0x08)[22]
0123456789ABCDEF
0xNUL1234³⁄5 SP 67¹89⁵⁄0
1xLTRS.⁹⁄:⁷⁄²?'DEL()=-/£+

Original Baudot, Continental European

[edit]
Original Baudot code, continental European variant (letter set, switched to with 0x10)[22]
0123456789ABCDEF
0xNULAEÉYUIOFIGSJGHBCFD
1x SP XZSTWVDELKMLRQNP
Original Baudot code, continental variant (figure set, switched to with 0x08)[22]
0123456789ABCDEF
0xNUL12&34º5 SP 67ʰ̵89ᶠ̵0
1xLTRS.,:;!?'DEL()=-/%

Original Baudot, ITA 1

[edit]
ITA 1 (letter set, switched to with 0x10)[22]
0123456789ABCDEF
0xNULAE CR YUIOFIGSJGHBCFD
1x SP  LF XZSTWVDELKMLRQNP
ITA 1 (figure set, switched to with 0x08)[22]
0123456789ABCDEF
0xNUL12 CR 34PU[a]5 SP 67+89PU[a]0
1xLTRS LF ,:.PU[a]?'DEL()=-/PU[a]%

Baudot–Murray variants

[edit]

Murray Code

[edit]
Murray code (letter set, switched to with 0x04)[22]
0123456789ABCDEF
0x SP ECOLALTRSSIU LF DRJNFCK
1xTZLWHYPQOBGFIGSMXVDEL/*[b]
Murray code (figure set, switched to with 0x1B)
0123456789ABCDEF
0x SP 3COLLTRS'87 LF ²4⁷⁄(⁹⁄
1x5./2⁵⁄6019?³⁄FIGS,£)DEL/*[b]

ITA 2 and US-TTY

[edit]
ITA2 and US-TTY Baudot–Murray code (letter set, switched to with 0x1F)
0123456789ABCDEF
0xNULE LF A SP SIU CR DRJNFCK
1xTZLWHYPQOBGFIGSMXVLTRS/DEL
US-TTY Baudot–Murray code (figure set, switched to with 0x1B)
0123456789ABCDEF
0xNUL3 LF  SP BEL87 CR $4',!:(
1x5")2#6019?&FIGS./;LTRS
ITA2 Baudot–Murray code (figure set, switched to with 0x1B)
0123456789ABCDEF
0xNUL3 LF  SP '87 CR ENQ4BEL,!:(
1x5+)2£6019?&FIGS./=LTRS

Weather code

[edit]

Meteorologists used a variant of ITA2 with the figures-case symbols, except for the ten digits, BEL and a few other characters, replaced by weather symbols:

Weather teleprinter encoding
Meteorological Baudot–Murray code (figure set, switched to with 0x1B)
0123456789ABCDEF
0x-3 LF  SP BEL87 CR 4
1x5+26019FIGS./LTRS

Details

[edit]
This sectionneeds additional citations forverification. Please helpimprove this article byadding citations to reliable sources in this section. Unsourced material may be challenged and removed.(November 2023) (Learn how and when to remove this message)

Note: This table presumes the space called "1" by Baudot and Murray is rightmost, and least significant. The way the transmitted bits were packed into larger codes varied by manufacturer. The most common solution allocates the bits from the least significant bit towards the most significant bit (leaving the three most significant bits of a byte unused).

Table of ITA2 codes (expressed ashexadecimal numbers)

In ITA2, characters are expressed using five bits. ITA2 uses two code sub-sets, the "letter shift" (LTRS), and the "figure shift" (FIGS). The FIGS character (11011) signals that the following characters are to be interpreted as being in the FIGS set, until this is reset by the LTRS (11111) character.[23] In use, the LTRS or FIGS shift key is pressed and released, transmitting the corresponding shift character to the other machine. The desired letters or figures characters are then typed. Unlike a typewriter or modern computer keyboard, the shift key isn't kept depressed whilst the corresponding characters are typed. "ENQuiry" will trigger the other machine's answerback. It means "Who are you?"

CR iscarriage return, LF isline feed, BEL is thebell character which rang a smallbell (often used to alert operators to an incoming message), SP is space, and NUL is thenull character (blank tape).

Note: the binary conversions of the codepoints are often shown in reverse order, depending on (presumably) from which side one views the paper tape. Note further that the"control" characters were chosen so that they were either symmetric or in useful pairs so that inserting a tape "upside down" did not result in problems for the equipment and the resulting printout could be deciphered. Thus FIGS (11011), LTRS (11111) and space (00100) are invariant, while CR (00010) and LF (01000), generally used as a pair, are treated the same regardless of order by page printers.[24] LTRS could also be used to overpunch characters to be deleted on a paper tape (much like DEL in 7-bitASCII).

The sequenceRYRYRY... is often used in test messages, and at the start of every transmission. Since R is 01010 and Y is 10101, the sequence exercises much of a teleprinter's mechanical components at maximum stress. Also, at one time, fine-tuning of the receiver was done using two coloured lights (one for each tone). 'RYRYRY...' produced 0101010101..., which made the lights glow with equal brightness when the tuning was correct. This tuning sequence is only useful when ITA2 is used with two-toneFSK modulation, such as is commonly seen inradioteletype (RTTY) usage.

US implementations of Baudot code may differ in the addition of a few characters, such as #, & on the FIGS layer.

The Russian version of Baudot code (MTK-2) used three shift modes; theCyrillic letter mode was activated by the character (00000). Because of the larger number of characters in the Cyrillic alphabet, the characters!,&,£ were omitted and replaced by Cyrillics, andBEL has the same code as Cyrillic letter Ю. The Cyrillic lettersЪ andЁ are omitted, and Ч is merged with the numeral 4.

See also

[edit]

Explanatory notes

[edit]
  1. ^abcd"At the disposal of each administration for its internal service"[22]
  2. ^ab"[G]ives invisible correction on page printers &* on slip printers."[22]

References

[edit]
  1. ^Ralston, Anthony; Reilly, Edwin D., eds. (1993), "Baudot Code",Encyclopedia of Computer Science (Third ed.), New York: IEEE Press/Van Nostrand Reinhold,ISBN 0-442-27679-6
  2. ^inRBK order
  3. ^abH. A. Emmons (1 May 1916)."Printer Systems".Wire & Radio Communications.34: 209.
  4. ^abcFischer, Eric N. (20 June 2000)."The Evolution of Character Codes, 1874–1968". ark:/13960/t07x23w8s. Retrieved20 December 2020.[...] In 1872, [Baudot] started research toward a telegraph system that would allow multiple operators to transmit simultaneously over a single wire and, as the transmissions were received, would print them in ordinary alphabetic characters on a strip of paper. He received a patent for such a system on June 17, 1874. [...] Instead of a variable delay followed by a single-unit pulse, Baudot's system used a uniform six time units to transmit each character. [...] his early telegraph probably used the six-unit code [...] that he attributes toDavy in an 1877 article. [...] in 1876 Baudot redesigned his equipment to use a five-unit code. Punctuation and digits were still sometimes needed, though, so he adopted fromHughes the use of two special letter space and figure space characters that would cause the printer to shift between cases at the same time as it advanced the paper without printing. The five-unit code he began using at this time [...] was structured to suit his keyboard [...], which controlled two units of each character with switches operated by the left hand and the other three units with the right hand. [...][1][2]
  5. ^Baudot, Jean-Maurice-Émile (June 1874)."Système de télégraphie rapide" (in French). ArchivesInstitut National de la Propriété Industrielle (INPI). Patent Brevet 103,898. Archived fromthe original on 16 December 2017.
  6. ^William V. Vansize (25 January 1901)."A New Page-Printing Telegraph".Transactions.18. American Institute of Electrical Engineers: 22.
  7. ^Procès d'Amiens Baudot vs Mimault
  8. ^abJennings, Tom (2020)."An annotated history of some character codes: Baudot's code".
  9. ^Beauchamp, K.G. (2001).History of Telegraphy: Its Technology and Application.Institution of Engineering and Technology. pp. 394–395.ISBN 0-85296-792-6.
  10. ^Alan G. Hobbs,5 Unit Codes, sectionBaudot Multiplex System
  11. ^Gleick, James (2011).The Information: A History, a Theory, a Flood. London: Fourth Estate. p. 203.ISBN 978-0-00-742311-8.
  12. ^Foster, Maximilian (August 1901)."A Successful Printing Telegraph".The World's Work: A History of Our Time.II:1195–1199. Retrieved9 July 2009.
  13. ^Copeland 2006, p. 38
  14. ^Telegraph and Telephone Age. 1921.I allocated the most frequently used letters in English language to the signals represented by the fewest holes in the perforated tape, and so on in proportion.
  15. ^"Telegraph Regulations and Final Protocol (Madrid, 1932)"(PDF). Archived fromthe original on 21 August 2023. Retrieved10 May 2024.
  16. ^Smith, Gil (2001)."Teletype Communication Codes"(PDF). Baudot.net.Archived(PDF) from the original on 20 August 2008. Retrieved11 July 2008.
  17. ^abSteinbuch, Karl W.; Weber, Wolfgang, eds. (1974) [1967].Taschenbuch der Informatik - Band III - Anwendungen und spezielle Systeme der Nachrichtenverarbeitung (in German). Vol. 3 (3 ed.). Berlin, Germany:Springer Verlag. pp. 328–329.ISBN 3-540-06242-4.LCCN 73-80607.{{cite book}}:|work= ignored (help)
  18. ^dataIP Limited."The "Baudot" Code". Archived fromthe original on 23 December 2017. Retrieved16 July 2017.
  19. ^ITU-TRecommendation S.2 / 11/1988, published in Fascicle VII.1 of theBlue Book
  20. ^"Enhanced Broadcast Solution – Interface Specification Final Version"(PDF). Deutsche Börse. 17 May 2010. Archived fromthe original(PDF) on 8 February 2012. Retrieved10 August 2011.
  21. ^Gillam, Richard (2002).Unicode Demystified. Addison-Wesley. p. 30.ISBN 0-201-70052-2.
  22. ^abcdefghi"Five-unit codes". NADCOMM museum. Archived fromthe original on 4 November 1999. Retrieved5 December 2001.
  23. ^This article is based on material taken fromBaudot+code at theFree On-line Dictionary of Computing prior to 1 November 2008 and incorporated under the "relicensing" terms of theGFDL, version 1.3 or later.
  24. ^Jennings, Tom (5 February 2020)."An annotated history of some character codes: ITA2". Retrieved1 June 2022.[...] the characters that are 'transmission control' related [...] are bit-wise symmetrical – the codes for FIGS, LTRS, space and BLANK – are the same reversed left to right! Further, the codes for CR and LF, equal each other when reversed left to right!
  25. ^Bacon, Francis (1605).The Proficience and Advancement of Learning Divine and Humane.

Further reading

[edit]

External links

[edit]
Early telecommunications
ISO/IEC 8859
Bibliographic use
National standards
ISO/IEC 2022
Mac OSCode pages
("scripts")
DOS code pages
IBM AIX code pages
Windows code pages
EBCDIC code pages
DEC terminals (VTx)
Platform specific
Unicode /ISO/IEC 10646
TeX typesetting system
Miscellaneous code pages
Control character
Related topics
Retrieved from "https://en.wikipedia.org/w/index.php?title=Baudot_code&oldid=1282238217"
Categories:
Hidden categories:

[8]ページ先頭

©2009-2025 Movatter.jp