Movatterモバイル変換

[0]ホーム

Jump to content

ASCII

Edit links

From Wikipedia, the free encyclopedia

(Redirected fromUS-ASCII)

Character encoding standard

This article is about the 7-bit character encoding standard. For other uses, seeASCII (disambiguation).Not to be confused with 8-bitextended ASCIIs.

ASCII
ASCII chart fromMIL-STD-188-100 (1972)
MIME / IANA	us-ascii
Alias(es)	ISO-IR-006,^[1] ANSI_X3.4-1968, ANSI_X3.4-1986, ISO_646.irv:1991, ISO646-US, us, IBM367, cp367^[2]
Languages	primarilyEnglish; also supportsMalay,Rotokas,Interlingua,Ido, andX-SAMPA
Classification	ISO/IEC 646 series
Extensions	Unicode ISO/IEC 8859 (series) KOI-8 OEM (series) Windows-125x (series) Others
Preceded by	ITA 2,FIELDATA
Succeeded by	ISO/IEC 8859,ISO/IEC 10646 (Unicode)

ASCII (/ˈæski/ ^ⓘASS-kee),^[3]^: 6 an acronym forAmerican Standard Code for Information Interchange, is acharacter encoding standard for representing a particular set of 95 (English language focused)printable and 33control characters – a total of 128code points. The set of available punctuation had significant impact on the syntax of computer languages and text markup. ASCII hugely influenced the design of character sets used by modern computers; for example, the first 128 code points ofUnicode are the same as ASCII.

ASCII encodes each code-point as a value from 0 to 127 – storable as a seven-bit integer.^[4] Ninety-five code-points are printable, including digits0 to9, lowercase lettersa toz, uppercase lettersA toZ, and commonly usedpunctuation symbols. For example, the letteri is represented as 105 (decimal). Also, ASCII specifies 33 non-printingcontrol codes which originated withTeletype devices, most of which are now obsolete.^[5] The control characters that are still commonly used includecarriage return,line feed, andtab.

ASCII lacks code-points for characters withdiacritical marks and therefore does not directly supportterms or names such asrésumé,jalapeño, orRené. But, depending on hardware and software support, some diacritical marks can berendered by overwriting a letter with abacktick (`) ortilde (~).

TheInternet Assigned Numbers Authority (IANA) prefers the nameUS-ASCII for this character encoding.^[2]

ASCII is one of theIEEE milestones.^[6]

	0	1	2	3	4	5	6	7	8	9	A	B	C	D	E	F
0x	NUL	SOH	STX	ETX	EOT	ENQ	ACK	BEL	BS	HT	LF	VT	FF	CR	SO	SI
1x	DLE	DC1	DC2	DC3	DC4	NAK	SYN	ETB	CAN	EM	SUB	ESC	FS	GS	RS	US
2x	SP	!	"	#	$	%	&	'	(	)	*	+	,	-	.	/
3x	0	1	2	3	4	5	6	7	8	9	:	;	<	=	>	?
4x	@	A	B	C	D	E	F	G	H	I	J	K	L	M	N	O
5x	P	Q	R	S	T	U	V	W	X	Y	Z	[	\	]	^	_
6x	`	a	b	c	d	e	f	g	h	i	j	k	l	m	n	o
7x	p	q	r	s	t	u	v	w	x	y	z	{	\|	}	~	DEL
Changed or added in 1963 version Changed in both 1963 version and 1965 draft

Binary	Oct	Dec	Hex	Abbreviation			UnicodeControl Pictures^[b]	Caret notation^[c]	C escape sequence^[d]	Name (1967)
Binary	Oct	Dec	Hex	1963	1965	1967	UnicodeControl Pictures^[b]	Caret notation^[c]	C escape sequence^[d]	Name (1967)
000 0000	000	0	00	NULL	NUL		␀	^@	\0^[e]	Null
000 0001	001	1	01	SOM	SOH		␁	^A		Start of Heading
000 0010	002	2	02	EOA	STX		␂	^B		Start of Text
000 0011	003	3	03	EOM	ETX		␃	^C		End of Text
000 0100	004	4	04	EOT			␄	^D		End of Transmission
000 0101	005	5	05	WRU	ENQ		␅	^E		Enquiry
000 0110	006	6	06	RU	ACK		␆	^F		Acknowledgement
000 0111	007	7	07	BELL	BEL		␇	^G	\a	Bell (Alert)
000 1000	010	8	08	FE0	BS		␈	^H	\b	Backspace^[f]^[g]
000 1001	011	9	09	HT/SK	HT		␉	^I	\t	Horizontal Tab^[h]
000 1010	012	10	0A	LF			␊	^J	\n	Line Feed
000 1011	013	11	0B	VTAB	VT		␋	^K	\v	Vertical Tab
000 1100	014	12	0C	FF			␌	^L	\f	Form Feed
000 1101	015	13	0D	CR			␍	^M	\r	Carriage Return^[i]
000 1110	016	14	0E	SO			␎	^N		Shift Out
000 1111	017	15	0F	SI			␏	^O		Shift In
001 0000	020	16	10	DC0	DLE		␐	^P		Data Link Escape
001 0001	021	17	11	DC1			␑	^Q		Device Control 1 (oftenXON)
001 0010	022	18	12	DC2			␒	^R		Device Control 2
001 0011	023	19	13	DC3			␓	^S		Device Control 3 (oftenXOFF)
001 0100	024	20	14	DC4			␔	^T		Device Control 4
001 0101	025	21	15	ERR	NAK		␕	^U		Negative Acknowledgement
001 0110	026	22	16	SYNC	SYN		␖	^V		Synchronous Idle
001 0111	027	23	17	LEM	ETB		␗	^W		End of Transmission Block
001 1000	030	24	18	S0	CAN		␘	^X		Cancel
001 1001	031	25	19	S1	EM		␙	^Y		End of Medium
001 1010	032	26	1A	S2	SS	SUB	␚	^Z		Substitute
001 1011	033	27	1B	S3	ESC		␛	^[	\e^[j]	Escape^[k]
001 1100	034	28	1C	S4	FS		␜	^\		File Separator
001 1101	035	29	1D	S5	GS		␝	^]		Group Separator
001 1110	036	30	1E	S6	RS		␞	^^^[l]		Record Separator
001 1111	037	31	1F	S7	US		␟	^_		Unit Separator
111 1111	177	127	7F	DEL			␡	^?		Delete^[m]^[g]

v t e Character encodings
Early telecommunications	Telegraph code Needle Morse Non-Latin Wabun/Kana Chinese Cyrillic Baudot and Murray Fieldata ASCII ISO/IEC 646 BCDIC Teletex andVideotex/Teletext T.51/ISO/IEC 6937 ITU T.61 ITU T.101 World System Teletext background sets Transcode
ISO/IEC 8859	Approved parts -1 (Western Europe) -2 (Central Europe) -3 (Maltese/Esperanto) -4 (North Europe) -5 (Cyrillic) -6 (Arabic) -7 (Greek) -8 (Hebrew) -9 (Turkish) -10 (Nordic) -11 (Thai) -13 (Baltic) -14 (Celtic) -15 (New Western Europe) -16 (Romanian) Abandoned parts -12 (Devanagari) Proposed but not approved KOI-8 Cyrillic Sámi Adaptations Welsh Estonian Ukrainian Cyrillic
Bibliographic use	MARC-8 ANSEL CCCII/EACC ISO 5426 5426-2 5427 5428 6438 6862
National standards	ArmSCII Big5 BraSCII BSCII CNS 11643 DIN 66003 ELOT 927 GOST 10859 GB 2312 GB 12345 GB 12052 GB 18030 HKSCS ISCII JIS X 0201 JIS X 0208 JIS X 0212 JIS X 0213 KOI-7 KPS 9566 KS X 1001 KS X 1002 LST 1564 LST 1590-4 PASCII Shift JIS SI 960 TIS-620 TSCII VISCII VSCII YUSCII
ISO/IEC 2022	ISO/IEC 8859 ISO/IEC 10367 Extended Unix Code / EUC
Mac OSCode pages ("scripts")	Armenian Arabic Barents Cyrillic Celtic Central European Croatian Cyrillic Devanagari Font X (Kermit) Gaelic Georgian Greek Gujarati Gurmukhi Hebrew Iceland Inuit Keyboard Latin (Kermit) Maltese/Esperanto Ogham Roman Romanian Sámi Turkish Turkic Cyrillic Ukrainian VT100
DOS code pages	437 737 850 858 861 862 863 864 865 866 867 868 869 899 904 932 936 942 949 950 951 1040 1043 1046 1098 1115 1116 1117 1118 1127 ABICOMP CS Indic CSX Indic CSX+ Indic CWI-2 Iran System Kamenický Mazovia MIK
IBM AIX code pages	895 896 912 915 921 922 1006 1008 1009 1010 1012 1013 1014 1015 1016 1017 1018 1019 1046 1133
Windows code pages	CER-GS 932 936 (GBK) 950 Extended Latin-8 1250 1251 1252 1253 1254 1255 1256 1257 1258 1270 Cyrillic + French Cyrillic + German Polytonic Greek
EBCDIC code pages	Japanese language in EBCDIC DKOI
DEC terminals (VTx)	Multinational (MCS) National Replacement (NRCS) French Canadian Swiss Spanish United Kingdom Dutch Finnish French Norwegian and Danish Swedish Norwegian and Danish (alternative) 8-bit Greek 8-bit Turkish SI 960 Hebrew Special Graphics Technical (TCS)
Platform specific	1052 1053 1054 1055 1058 Acorn RISC OS Amstrad CPC Apple II ATASCII Atari ST BICS Casio calculators CDC Compucolor 8001 Compucolor II CP/M+ DEC RADIX 50 DEC MCS/NRCS DG International Galaksija GEM GSM 03.38 HP Roman HP FOCAL HP RPL SQUOZE LICS LMBCS MSX NEC APC NeXT PETSCII PostScript Standard PostScript Latin 1 SAM Coupé Sega SC-3000 Sharp calculators Sharp MZ Sinclair QL Teletext TI calculators TRS-80 Ventura International WISCII XCCS ZX80 ZX81 ZX Spectrum
Unicode /ISO/IEC 10646	UTF-1 UTF-7 UTF-8 UTF-16 UTF-32 UTF-EBCDIC GB 18030 DIN 91379 BOCU-1 CESU-8 SCSU TACE16 Comparison of Unicode encodings
TeX typesetting system	Cork LY1 OML OMS OT1
Miscellaneous code pages	ABICOMP ASMO 449 Digital encoding of APL symbols ISO-IR-68 ARIB STD-B24 Fieldata HZ IEC-P27-1 INIS 7-bit 8-bit ISO-IR-169 ISO 2033 KOI KOI8-R KOI8-RU KOI8-U Mojikyō SEASCII Stanford/ITS Symbol TRON Unified Hangul Code
Control character	Morse prosigns C0 and C1 control codes ISO/IEC 6429 JIS X 0211 Unicode control, format and separator characters Whitespace characters
Related topics	CCSID Character encodings in HTML Charset detection Han unification Hardware code page MICR code Mojibake Variable-length encoding
Character sets

Binary	Oct	Dec	Hex	Glyph
Binary	Oct	Dec	Hex	1963	1965	1967
010 0000	040	32	20	space (no visible glyph)
010 0001	041	33	21	!
010 0010	042	34	22	"
010 0011	043	35	23	#
010 0100	044	36	24	$
010 0101	045	37	25	%
010 0110	046	38	26	&
010 0111	047	39	27	'
010 1000	050	40	28	(
010 1001	051	41	29	)
010 1010	052	42	2A	*
010 1011	053	43	2B	+
010 1100	054	44	2C	,
010 1101	055	45	2D	-
010 1110	056	46	2E	.
010 1111	057	47	2F	/
011 0000	060	48	30	0
011 0001	061	49	31	1
011 0010	062	50	32	2
011 0011	063	51	33	3
011 0100	064	52	34	4
011 0101	065	53	35	5
011 0110	066	54	36	6
011 0111	067	55	37	7
011 1000	070	56	38	8
011 1001	071	57	39	9
011 1010	072	58	3A	:
011 1011	073	59	3B	;
011 1100	074	60	3C	<
011 1101	075	61	3D	=
011 1110	076	62	3E	>
011 1111	077	63	3F	?
100 0000	100	64	40	@	`	@
100 0001	101	65	41	A
100 0010	102	66	42	B
100 0011	103	67	43	C
100 0100	104	68	44	D
100 0101	105	69	45	E
100 0110	106	70	46	F

Movatterモバイル変換

History

Revisions

Design considerations

Bit width

Internal organization

Character order

Character set

Character groups

Control characters

Delete vs backspace

Escape

End of line

End of file/stream

Table of codes

Control code table

Printable character table

Usage

Variants and derivations

7-bit codes

8-bit codes

Unicode

See also

Notes

References

Further reading

External links