PDF Reference sixth edition, Adobe Portable Document Format Version 1.7 (book 2) — page 11

1003
SECTION D.2
PDFDocEncoding Character Set
CHAR-
DEC
HEX
OCTAL
UNICODE
UNICODE CHARACTER NAME OR (ALTERNATIVE
NOTES
ACTER
ALIAS)
:
58
0x3a
0072
U+003A
COLON
;
59
0x3b
0073
U+003B
SEMICOLON
<
60
0x3c
0074
U+003C
LESS THAN SIGN (<)
SR
=
61
0x3d
0075
U+003D
EQUALS SIGN
>
62
0x3e
0076
U+003E
GREATER THAN SIGN (>)
?
63
0x3f
0077
U+003F
QUESTION MARK
@
64
0x40
0100
U+0040
COMMERCIAL AT
A
65
0x41
0101
U+0041
B
66
0x42
0102
U+0042
C
67
0x43
0103
U+0043
D
68
0x44
0104
U+0044
E
69
0x45
0105
U+0045
F
70
0x46
0106
U+0046
G
71
0x47
0107
U+0047
H
72
0x48
0110
U+0048
I
73
0x49
0111
U+0049
J
74
0x4a
0112
U+004A
K
75
0x4b
0113
U+004B
L
76
0x4c
0114
U+004C
M
77
0x4d
0115
U+004D
N
78
0x4e
0116
U+004E
O
79
0x4f
0117
U+004F
P
80
0x50
0120
U+0050
Q
81
0x51
0121
U+0051
R
82
0x52
0122
U+0052
S
83
0x53
0123
U+0053
T
84
0x54
0124
U+0054
U
85
0x55
0125
U+0055
V
86
0x56
0126
U+0056
W
87
0x57
0127
U+0057
X
88
0x58
0130
U+0058
Y
89
0x59
0131
U+0059
1004
APPENDIX D
Character Sets and Encodings
CHAR-
DEC
HEX
OCTAL
UNICODE
UNICODE CHARACTER NAME OR (ALTERNATIVE
NOTES
ACTER
ALIAS)
Z
90
0x5a
0132
U+005A
[
91
0x5b
0133
U+005B
LEFT SQUARE BRACKET
\
92
0x5c
0134
U+005C
REVERSE SOLIDUS (backslash)
]
93
0x5d
0135
U+005D
RIGHT SQUARE BRACKET
^
94
0x5e
0136
U+005E
CIRCUMFLEX ACCENT (hat)
_
95
0x5f
0137
U+005F
LOW LINE (SPACING UNDERSCORE)
`
96
0x60
0140
U+0060
GRAVE ACCENT
a
97
0x61
0141
U+0061
b
98
0x62
0142
U+0062
c
99
0x63
0143
U+0063
d
100
0x64
0144
U+0064
e
101
0x65
0145
U+0065
f
102
0x66
0146
U+0066
g
103
0x67
0147
U+0067
h
104
0x68
0150
U+0068
i
105
0x69
0151
U+0069
j
106
0x6a
0152
U+006A
k
107
0x6b
0153
U+006B
l
108
0x6c
0154
U+006C
m
109
0x6d
0155
U+006D
n
110
0x6e
0156
U+006E
o
111
0x6f
0157
U+006F
p
112
0x70
0160
U+0070
q
113
0x71
0161
U+0071
r
114
0x72
0162
U+0072
s
115
0x73
0163
U+0073
t
116
0x74
0164
U+0074
u
117
0x75
0165
U+0075
v
118
0x76
0166
U+0076
w
119
0x77
0167
U+0077
x
120
0x78
0170
U+0078
y
121
0x79
0171
U+0079
1005
SECTION D.2
PDFDocEncoding Character Set
CHAR-
DEC
HEX
OCTAL
UNICODE
UNICODE CHARACTER NAME OR (ALTERNATIVE
NOTES
ACTER
ALIAS)
z
122
0x7a
0172
U+007A
{
123
0x7b
0173
U+007B
LEFT CURLY BRACKET
|
124
0x7c
0174
U+007C
VERTICAL LINE
}
125
0x7d
0175
U+007D
RIGHT CURLY BRACKET
~
126
0x7e
0176
U+007E
TILDE
127
0x7f
0177
Undefined
U
128
0x80
0200
U+2022
BULLET
129
0x81
0201
U+2020
DAGGER
130
0x82
0202
U+2021
DOUBLE DAGGER
131
0x83
0203
U+2026
HORIZONTAL ELLIPSIS
132
0x84
0204
U+2014
EM DASH
133
0x85
0205
U+2013
EN DASH
ƒ
134
0x86
0206
U+0192
135
0x87
0207
U+2044
FRACTION SLASH (solidus)
136
0x88
0210
U+2039
SINGLE LEFT-POINTING ANGLE
QUOTATION MARK
137
0x89
0211
U+203A
SINGLE RIGHT-POINTING ANGLE
QUOTATION MARK
Š
138
0x8a
0212
U+2212
139
0x8b
0213
U+2030
PER MILLE SIGN
140
0x8c
0214
U+201E
DOUBLE LOW-9 QUOTATION MARK
(quotedblbase)
141
0x8d
0215
U+201C
LEFT DOUBLE QUOTATION MARK (double
quote left)
142
0x8e
0216
U+201D
RIGHT DOUBLE QUOTATION MARK
(quotedblright)
143
0x8f
0217
U+2018
LEFT SINGLE QUOTATION MARK (quoteleft)
144
0x90
0220
U+2019
RIGHT SINGLE QUOTATION MARK
(quoteright)
145
0x91
0221
U+201A
SINGLE LOW-9 QUOTATION MARK
(quotesinglbase)
146
0x92
0222
U+2122
TRADE MARK SIGN
fi
147
0x93
0223
U+FB01
LATIN SMALL LIGATURE FI
1006
APPENDIX D
Character Sets and Encodings
CHAR-
DEC
HEX
OCTAL
UNICODE
UNICODE CHARACTER NAME OR (ALTERNATIVE
NOTES
ACTER
ALIAS)
fl
148
0x94
0224
U+FB02
LATIN SMALL LIGATURE FL
149
0x95
0225
U+0141
LATIN CAPITAL LETTER L WITH STROKE
OE
150
0x96
0226
U+0152
LATIN CAPITAL LIGATURE OE
Š
151
0x97
0227
U+0160
LATIN CAPITAL LETTER S WITH CARON
Ÿ
152
0x98
0230
U+0178
LATIN CAPITAL LETTER Y WITH DIAERESIS
Z hat
153
0x99
0231
U+017D
LATIN CAPITAL LETTER Z WITH CARON
i
154
0x9a
0232
U+0131
LATIN SMALL LETTER DOTLESS I
l/
155
0x9b
0233
U+0142
LATIN SMALL LETTER L WITH STROKE
œ
156
0x9c
0234
U+0153
LATIN SMALL LIGATURE OE
š
157
0x9d
0235
U+0161
LATIN SMALL LETTER S WITH CARON
ž
158
0x9e
0236
U+017E
LATIN SMALL LETTER Z WITH CARON
159
0x9f
0237
Undefined
U
160
0xa0
0240
U+20AC
EURO SIGN
¡
161
0xa1
0241
U+00A1
INVERTED EXCLAMATION MARK
¢
162
0xa2
0242
U+00A2
CENT SIGN
£
163
0xa3
0243
U+00A3
POUND SIGN (sterling)
¤
164
0xa4
0244
U+00A4
CURRENCY SIGN
¥
165
0xa5
0245
U+00A5
YEN SIGN
¦
166
0xa6
0246
U+00A6
BROKEN BAR
§
167
0xa7
0247
U+00A7
SECTION SIGN
¨
168
0xa8
0250
U+00A8
DIAERESIS
169
0xa9
0251
U+00A9
COPYRIGHT SIGN
ª
170
0xaa
0252
U+00AA
FEMININE ORDINAL INDICATOR
«
171
0xab
0253
U+00AB
LEFT-POINTING DOUBLE ANGLE
QUOTATION MARK
¬
172
0xac
0254
U+00AC
NOT SIGN
173
0xad
0255
Undefined
U
®
174
0xae
0256
U+00AE
REGISTERED SIGN
¯
175
0xaf
0257
U+00AF
MACRON
°
176
0xb0
0260
U+00B0
DEGREE SIGN
±
177
0xb1
0261
U+00B1
PLUS-MINUS SIGN
2
178
0xb2
0262
U+00B2
SUPERSCRIPT TWO
1007
SECTION D.2
PDFDocEncoding Character Set
CHAR-
DEC
HEX
OCTAL
UNICODE
UNICODE CHARACTER NAME OR (ALTERNATIVE
NOTES
ACTER
ALIAS)
3
179
0xb3
0263
U+00B3
SUPERSCRIPT THREE
´
180
0xb4
0264
U+00B4
ACUTE ACCENT
μ
181
0xb5
0265
U+00B5
MICRO SIGN
182
0xb6
0266
U+00B6
PILCROW SIGN
·
183
0xb7
0267
U+00B7
MIDDLE DOT
¸
184
0xb8
0270
U+00B8
CEDILLA
1
185
0xb9
0271
U+00B9
SUPERSCRIPT ONE
º
186
0xba
0272
U+00BA
MASCULINE ORDINAL INDICATOR
»
187
0xbb
0273
U+00BB
RIGHT-POINTING DOUBLE ANGLE
QUOTATION MARK
¼
188
0xbc
0274
U+00BC
VULGAR FRACTION ONE QUARTER
½
189
0xbd
0275
U+00BD
VULGAR FRACTION ONE HALF
¾
190
0xbe
0276
U+00BE
VULGAR FRACTION THREE QUARTERS
¿
191
0xbf
0277
U+00BF
INVERTED QUESTION MARK
À
192
0xc0
0300
U+00C0
Á
193
0xc1
0301
U+00C1
Â
194
0xc2
0302
U+00C2
Ã
195
0xc3
0303
U+00C3
Ä
196
0xc4
0304
U+00C4
Å
197
0xc5
0305
U+00C5
Æ
198
0xc6
0306
U+00C6
Ç
199
0xc7
0307
U+00C7
È
200
0xc8
0310
U+00C8
É
201
0xc9
0311
U+00C9
Ê
202
0xca
0312
U+00CA
Ë
203
0xcb
0313
U+00CB
Ì
204
0xcc
0314
U+00CC
Í
205
0xcd
0315
U+00CD
Î
206
0xce
0316
U+00CE
Ï
207
0xcf
0317
U+00CF
Ð
208
0xd0
0320
U+00D0
Ñ
209
0xd1
0321
U+00D1
1008
APPENDIX D
Character Sets and Encodings
CHAR-
DEC
HEX
OCTAL
UNICODE
UNICODE CHARACTER NAME OR (ALTERNATIVE
NOTES
ACTER
ALIAS)
Ò
210
0xd2
0322
U+00D2
Ó
211
0xd3
0323
U+00D3
Ô
212
0xd4
0324
U+00D4
Õ
213
0xd5
0325
U+00D5
Ö
214
0xd6
0326
U+00D6
×
215
0xd7
0327
U+00D7
Ø
216
0xd8
0330
U+00D8
Ù
217
0xd9
0331
U+00D9
Ú
218
0xda
0332
U+00DA
Û
219
0xdb
0333
U+00DB
Ü
220
0xdc
0334
U+00DC
Ý
221
0xdd
0335
U+00DD
Þ
222
0xde
0336
U+00DE
ß
223
0xdf
0337
U+00DF
à
224
0xe0
0340
U+00E0
á
225
0xe1
0341
U+00E1
â
226
0xe2
0342
U+00E2
ã
227
0xe3
0343
U+00E3
ä
228
0xe4
0344
U+00E4
å
229
0xe5
0345
U+00E5
æ
230
0xe6
0346
U+00E6
ç
231
0xe7
0347
U+00E7
è
232
0xe8
0350
U+00E8
é
233
0xe9
0351
U+00E9
ê
234
0xea
0352
U+00EA
ë
235
0xeb
0353
U+00EB
ì
236
0xec
0354
U+00EC
í
237
0xed
0355
U+00ED
î
238
0xee
0356
U+00EE
ï
239
0xef
0357
U+00EF
ð
240
0xf0
0360
U+00F0
ñ
241
0xf1
0361
U+00F1
1009
SECTION D.2
PDFDocEncoding Character Set
CHAR-
DEC
HEX OCTAL
UNICODE UNICODE CHARACTER NAME OR (ALTERNATIVE
NOTES
ACTER
ALIAS)
ò
242
0xf2
0362
U+00F2
ó
243
0xf3
0363
U+00F3
ô
244
0xf4
0364
U+00F4
õ
245
0xf5
0365
U+00F5
ö
246
0xf6
0366
U+00F6
÷
247
0xf7
0367
U+00F7
ø
248
0xf8
0370
U+00F8
ù
249
0xf9
0371
U+00F9
ú
250
0xfa
0372
U+00FA
û
251
0xfb
0373
U+00FB
ü
252
0xfc
0374
U+00FC
ý
253
0xfd
0375
U+00FD
þ
254
0xfe
0376
U+00FE
ÿ
255
0xff
0377
U+00FF
1010
APPENDIX D
Character Sets and Encodings
D.3
Expert Set and MacExpertEncoding
CHAR
NAME
CODE
CHAR
NAME
CODE
æ
AEsmall
276
j
Jsmall
152
á
Aacutesmall
207
k
Ksmall
153
â
Acircumflexsmall
211
ł
Lslashsmall
302
´
Acutesmall
047
l
Lsmall
154
ä
Adieresissmall
212
¯
Macronsmall
364
à
Agravesmall
210
m
Msmall
155
å
Aringsmall
214
n
Nsmall
156
a
Asmall
141
ñ
Ntildesmall
226
ã
Atildesmall
213
œ
OEsmall
317
˘
Brevesmall
363
ó
Oacutesmall
227
b
Bsmall
142
ô
Ocircumflexsmall
231
ˇ
Caronsmall
256
ö
Odieresissmall
232
ç
Ccedillasmall
215
˛
Ogoneksmall
362
¸
Cedillasmall
311
ò
Ogravesmall
230
ˆ
Circumflexsmall
136
ø
Oslashsmall
277
c
Csmall
143
o
Osmall
157
¨
Dieresissmall
254
õ
Otildesmall
233
˙
Dotaccentsmall
372
p
Psmall
160
d
Dsmall
144
q
Qsmall
161
é
Eacutesmall
216
˚
Ringsmall
373
ê
Ecircumflexsmall
220
r
Rsmall
162
ë
Edieresissmall
221
š
Scaronsmall
247
è
Egravesmall
217
s
Ssmall
163
e
Esmall
145
þ
Thornsmall
271
ð
Ethsmall
104
˜
Tildesmall
176
f
Fsmall
146
t
Tsmall
164
`
Gravesmall
140
ú
Uacutesmall
234
g
Gsmall
147
û
Ucircumflexsmall
236
h
Hsmall
150
ü
Udieresissmall
237
˝
Hungarumlautsmall
042
ù
Ugravesmall
235
í
Iacutesmall
222
u
Usmall
165
î
Icircumflexsmall
224
v
Vsmall
166
ï
Idieresissmall
225
w
Wsmall
167
ì
Igravesmall
223
x
Xsmall
170
i
Ismall
151
ý
Yacutesmall
264
1011
SECTION D.3
Expert Set and MacExpertEncoding
CHAR
NAME
CODE
CHAR
NAME
CODE
ÿ
Ydieresissmall
330
4
fouroldstyle
064
y
Ysmall
171
foursuperior
335
ž
Zcaronsmall
275
fraction
057
z
Zsmall
172
-
hyphen
055
&
ampersandsmall
046
-
hypheninferior
137
a
asuperior
201
-
hyphensuperior
321
b
bsuperior
365
i
isuperior
351
¢
centinferior
251
l
lsuperior
361
¢
centoldstyle
043
m
msuperior
367
¢
centsuperior
202
nineinferior
273
:
colon
072
9
nineoldstyle
071
colonmonetary
173
ninesuperior
341
,
comma
054
nsuperior
366
,
commainferior
262
onedotenleader
053
,
commasuperior
370
oneeighth
112
$
dollarinferior
266
1
onefitted
174
$
dollaroldstyle
044
½
onehalf
110
$
dollarsuperior
045
oneinferior
301
d
dsuperior
353
1
oneoldstyle
061
eightinferior
245
¼
onequarter
107
8
eightoldstyle
070
¹
onesuperior
332
eightsuperior
241
onethird
116
e
esuperior
344
o
osuperior
257
¡
exclamdownsmall
326
parenleftinferior
133
!
exclamsmall
041
parenleftsuperior
050
ff
ff
126
parenrightinferior
135
ffi
ffi
131
parenrightsuperior
051
ffl
ffl
132
period
056
fi
fi
127
periodinferior
263
figuredash
320
periodsuperior
371
fiveeighths
114
¿
questiondownsmall
300
fiveinferior
260
?
questionsmall
077
5
fiveoldstyle
065
r
rsuperior
345
fivesuperior
336
R
rupiah
175
fl
fl
130
;
semicolon
073
fourinferior
242
seveneighths
115
1012
APPENDIX D
Character Sets and Encodings
CHAR
NAME
CODE
CHAR
NAME
CODE
seveninferior
246
threequartersemdash
075
7
sevenoldstyle
067
³
threesuperior
334
sevensuperior
340
t
tsuperior
346
sixinferior
244
twodotenleader
052
6
sixoldstyle
066
twoinferior
252
sixsuperior
337
2
twooldstyle
062
space
040
²
twosuperior
333
s
ssuperior
352
twothirds
117
threeeighths
113
zeroinferior
274
threeinferior
243
0
zerooldstyle
060
3
threeoldstyle
063
zerosuperior
342
¾
threequarters
111
1013
SECTION D.4
Symbol Set and Encoding
D.4
Symbol Set and Encoding
CHAR
NAME
CODE
CHAR
NAME
CODE
I
Alpha
101
C
arrowboth
253
J
Beta
102
‹
arrowdblboth
333
]
Chi
103
‡
arrowdbldown
337
6
Delta
104
ˆ
arrowdblleft
334
L
Epsilon
105
‰
arrowdblright
336
N
Eta
110
Š
arrowdblup
335
Euro
240
?
arrowdown
257
K
Gamma
107
±
arrowhorizex
276
¼
Ifraktur
301
@
arrowleft
254
P
Iota
111
A
arrowright
256
Q
Kappa
113
B
arrowup
255
R
Lambda
114
°
arrowvertex
275
S
Mu
115
›
asteriskmath
052
T
Nu
116
|
bar
174
1
Omega
127
`
beta
142
V
Omicron
117
{
braceleft
173
\
Phi
106
}
braceright
175
W
Pi
120
¨
bracelefttp
354
^
Psi
131
braceleftmid
355
½
Rfraktur
302
ª
braceleftbt
356
X
Rho
122
¬
bracerighttp
374
Y
Sigma
123
­
bracerightmid
375
Z
Tau
124
®
bracerightbt
376
O
Theta
121
«
braceex
357
[
Upsilon
125
[
bracketleft
133
¯
Upsilon1
241
]
bracketright
135
U
Xi
130
•
bracketlefttp
351
M
Zeta
132
³
bracketleftex
352
”
aleph
300
–
bracketleftbt
353
_
alpha
141
—
bracketrighttp
371
&
ampersand
046
µ
bracketrightex
372
angle
320
˜
bracketrightbt
373
angleleft
341
bullet
267
œ
angleright
361
¿
carriagereturn
277
5
approxequal
273
r
chi
143
1014
APPENDIX D
Character Sets and Encodings
CHAR
NAME
CODE
CHAR
NAME
CODE
€
circlemultiply
304
H
integralbt
365
circleplus
305
E
intersection
307
y
club
247
f
iota
151
:
colon
072
g
kappa
153
,
comma
054
h
lambda
154
congruent
100
<
less
074
copyrightsans
343
)
lessequal
243
copyrightserif
323
Ž
logicaland
331
°
degree
260
¬
logicalnot
330
b
delta
144
logicalor
332
z
diamond
250
9
lozenge
340
÷
divide
270
<
minus
055
u
dotmath
327
v
minute
242
8
eight
070
μ
mu
155
D
element
316
×
multiply
264
ellipsis
274
9
nine
071
’
emptyset
306
notelement
317
¡
epsilon
145
&
notequal
271
=
equal
075
†
notsubset
313
>
equivalence
272
i
nu
156
d
eta
150
#
numbersign
043
!
exclam
041
t
omega
167
š
existential
044
Ÿ
omega1
166
5
five
065
k
omicron
157
ƒ
florin
246
1
one
061
4
four
064
(
parenleft
050
fraction
244
)
parenright
051
a
gamma
147
£
parenlefttp
346
¢
gradient
321
²
parenleftex
347
>
greater
076
¤
parenleftbt
350
greaterequal
263
¥
parenrighttp
366
x
heart
251
´
parenrightex
367
'
infinity
245
¦
parenrightbt
370
0
integral
362
,
partialdiff
266
G
integraltp
363
%
percent
045
“
integralex
364
period
056
1015
SECTION D.4
Symbol Set and Encoding
CHAR
NAME
CODE
CHAR
NAME
CODE
Œ
perpendicular
136
¾
similar
176
q
phi
146
6
six
066
phi1
152
/
slash
057
/
pi
160
space
040
+
plus
053
{
spade
252
±
plusminus
261
~
suchthat
047
product
325
-
summation
345
„
propersubset
314
o
tau
164
‚
propersuperset
311
‘
therefore
134
|
proportional
265
e
theta
161
s
psi
171
ž
theta1
112
?
question
077
3
three
063
3
radical
326
trademarksans
344
radicalex
140
trademarkserif
324
reflexsubset
315
2
two
062
ƒ
reflexsuperset
312
_
underscore
137
®
registersans
342
F
union
310
®
registerserif
322
™
universal
042
l
rho
162
p
upsilon
165
w
second
262
§
weierstrass
303
;
semicolon
073
j
xi
170
7
seven
067
0
zero
060
m
sigma
163
c
zeta
172
n
sigma1
126
1016
APPENDIX D
Character Sets and Encodings
D.5
ZapfDingbats Set and Encoding
CHAR
NAME
CODE
CHAR
NAME
CODE
CHAR
NAME
CODE
CHAR
NAME
CODE
space
040
a30
103
a65
146
a109
253
a1
041
a31
104
a66
147
a120
254
a2
042
a32
105
a67
150
a121
255
a202
043
a33
106
a68
151
a122
256
a3
044
a34
107
a69
152
a123
257
a4
045
a35
110
a70
153
a124
260
a5
046
a36
111
a71
154
a125
261
a119
047
a37
112
a72
155
a126
262
a118
050
a38
113
a73
156
a127
263
a117
051
a39
114
a74
157
a128
264
a11
052
a40
115
a203
160
a129
265
a12
053
a41
116
a75
161
a130
266
a13
054
a42
117
a204
162
a131
267
a14
055
a43
120
a76
163
a132
270
a15
056
a44
121
a77
164
a133
271
a16
057
a45
122
a78
165
a134
272
a105
060
a46
123
a79
166
a135
273
a17
061
a47
124
a81
167
a136
274
a18
062
a48
125
a82
170
a137
275
a19
063
a49
126
a83
171
a138
276
a20
064
a50
127
a84
172
a139
277
a21
065
a51
130
a97
173
a140
300
a22
066
a52
131
a98
174
a141
301
a23
067
a53
132
a99
175
a142
302
a24
070
a54
133
a100
176
a143
303
a25
071
a55
134
a101
241
a144
304
a26
072
a56
135
a102
242
a145
305
a27
073
a57
136
a103
243
a146
306
a28
074
a58
137
a104
244
a147
307
a6
075
a59
140
a106
245
a148
310
a7
076
a60
141
a107
246
a149
311
a8
077
a61
142
a108
247
a150
312
a9
100
a62
143
a112
250
a151
313
a10
101
a63
144
a111
251
a152
314
a29
102
a64
145
a110
252
a153
315
1017
SECTION D.5
ZapfDingbats Set and Encoding
CHAR NAME CODE
CHAR NAME CODE
CHAR NAME CODE
CHAR NAME CODE
a154
316
a192
332
a176
346
a184
363
a155
317
a166
333
a177
347
a197
364
a156
320
a167
334
a178
350
a185
365
a157
321
a168
335
a179
351
a194
366
a158
322
a169
336
a193
352
a198
367
a159
323
a170
337
a180
353
a186
370
a160
324
a171
340
a199
354
a195
371
a161
325
a172
341
a181
355
a187
372
a163
326
a173
342
a200
356
a188
373
a164
327
a162
343
a182
357
a189
374
a196
330
a174
344
a201
361
a190
375
a165
331
a175
345
a183
362
a191
376
1018
APPENDIX D
Character Sets and Encodings
APPENDIX E
PDF Name Registry
E
This appendix discusses a registry, maintained by Adobe for developers, that con-
tains private names and formats used by PDF producers or Acrobat plug-in ex-
tensions.
Acrobat enables third parties to add private data to PDF documents and to add
plug-in extensions that change viewer behavior based on this data. However,
Acrobat users have certain expectations when opening a PDF document, no mat-
ter what plug-ins are available. PDF enforces certain restrictions on private data
in order to meet these expectations.
A PDF producer or Acrobat viewer plug-in extension may define new types of
actions, destinations, annotations, security, and file system handlers. If a user
opens a PDF document and the plug-in that implements the new type of object is
unavailable, the viewer behaves as described in Appendix H, “Compatibility and
Implementation Notes.”
A PDF producer or Acrobat plug-in extension may also add keys to any PDF
object that is implemented as a dictionary, except the file trailer dictionary (see
Section 3.4.4, “File Trailer”). In addition, a PDF producer or Acrobat plug-in may
create tags that indicate the role of marked-content operators (PDF 1.2), as
described in Section 10.5, “Marked Content.”
To avoid conflicts with third-party names and with future versions of PDF, Adobe
maintains a registry for certain private names and formats. Developers must only
add private data that conforms to the registry rules. The registry includes three
classes:
First class. Names and data formats that are of value to a wide range of de-
velopers. All names defined in any version of the PDF specification are first-
1019
1020
APPENDIX E
PDF Name Registry
class names. Plug-in extensions that are publicly available should often use
first-class names for their private data. First-class names and data formats must
be registered with Adobe and are made available for all developers to use. To
submit a private name and format for consideration as first-class, see the link
on registering a private PDF extension, at the following Web page:
Second class. Names that are applicable to a specific developer. (Adobe does not
register second-class data formats.) Adobe distributes second-class names by
registering developer-specific prefixes, which must be used as the first char-
acters in the names of all private data added by the developer. Adobe will not
register the same prefix to two different developers, thereby ensuring that dif-
ferent developers’ second-class names do not conflict. It is the responsibility of
the developer not to use the same name in conflicting ways. To register a devel-
oper-specific prefix, use the Acrobat SDK feedback form accessible through the
following Web page:
Third class. Names that can be used only in files that other third parties will
never see because they may conflict with third-class names defined by others.
Third-class names all begin with a specific prefix reserved by Adobe for private
plug-in extensions. This prefix, which is XX, must be used as the first characters
in the names of all private data added by the developer. It is not necessary to
contact Adobe to register third-class names.
Note: New keys for the document information dictionary (see Section 10.2.1, “Doc-
ument Information Dictionary”) or a thread information dictionary (in the I entry
of a thread dictionary; see Section 8.3.2, “Articles”) need not be registered.
APPENDIX F
Linearized PDF
F
A Linearized PDF file is a file that has been organized in a special way to enable
efficient incremental access in a network environment. The file is valid PDF in all
respects, and is compatible with all existing viewers and other PDF applications.
Enhanced viewer applications can recognize that a PDF file has been linearized
and can take advantage of that organization (as well as added hint information) to
enhance viewing performance.
The Linearized PDF file organization is an optional feature available beginning in
PDF 1.2. Its primary goal is to achieve the following behavior:
When a document is opened, display the first page as quickly as possible. The
first page to be viewed can be an arbitrary page of the document, not neces-
sarily page 0 (though opening at page 0 is most common).
When the user requests another page of an open document (for example, by
going to the next page or by following a link to an arbitrary page), display that
page as quickly as possible.
When data for a page is delivered over a slow channel, display the page incre-
mentally as it arrives. To the extent possible, display the most useful data first.
Permit user interaction, such as following a link, to be performed even before
the entire page has been received and displayed.
This behavior should be achieved for documents of arbitrary size. The total num-
ber of pages in the document should have little or no effect on the user-perceived
performance of viewing any particular page.
The primary focus of Linearized PDF is optimized viewing of read-only PDF
documents. It is intended that the Linearized PDF be generated once and read
many times. Incremental update is still permitted, but the resulting PDF is no
1021
1022
APPENDIX F
Linearized PDF
longer linearized and subsequently is treated as ordinary PDF. Linearizing it
again may require reprocessing the entire file; see Section F.4.6, “Accessing an Up-
dated File,” for details.
Linearized PDF requires two additions to the PDF specification:
Rules for the ordering of objects in the PDF file
Additional data structures, called hint tables, that enable efficient navigation
within the document
Both of these additions are relatively simple to describe; however, using them
effectively requires a deeper understanding of their purpose. Consequently, this
appendix goes considerably beyond a simple specification of these PDF exten-
sions to include background, motivation, and strategies.
Section F.1, “Background and Assumptions,” provides background information
about the properties of the Web that are relevant to the design of Linearized
PDF.
Section F.2, “Linearized PDF Document Structure,” specifies the file format and
object-ordering requirements of Linearized PDF.
Section F.3, “Hint Tables,” specifies the detailed representation of the hint
tables.
Section F.4, “Access Strategies,” outlines strategies for accessing Linearized PDF
over a network, which in turn determine the optimal way to organize the PDF
file.
The reader is assumed to be familiar with the basic architecture of the Web, in-
cluding terms such as URL, HTTP, and MIME.
F.1
Background and Assumptions
The principal problem addressed by the Linearized PDF design is the access of
PDF documents through the Web. This environment has the following important
properties:
The access protocol (HTTP) is a transaction consisting of a request and a re-
sponse. The client presents a request in the form of a URL, and the server sends
a response consisting of one or more MIME-tagged data blocks.
1023
SECTION F.1
Background and Assumptions
After a transaction has completed, obtaining more data requires a new request-
response transaction. The connection between client and server does not ordi-
narily persist beyond the end of a transaction, although some implementations
may attempt to cache the open connection to expedite subsequent transactions
with the same server.
Round-trip delay can be significant. A request-response transaction can take
up to several seconds, independent of the amount of data requested.
The data rate may be limited. A typical bottleneck is a slow modem link be-
tween the client and the Internet service provider.
These properties are generally shared by other wide-area network architectures
besides the Web. Also, CD-ROMs share some of these properties, since they have
relatively slow seek times and limited data rates compared to magnetic media.
The remainder of this appendix focuses on the Web.
Some additional properties of the HTTP protocol are relevant to the problem of
accessing PDF files efficiently. These properties may not all be shared by other
protocols or network environments.
When a PDF file is initially accessed (such as by following a URL hyperlink
from some other document), the file type is not known to the client. Therefore,
the client initiates a transaction to retrieve the entire document and then in-
spects the MIME tag of the response as it arrives. Only at that point is the doc-
ument known to be PDF. Additionally, with a properly configured server
environment, the length of the document becomes known at that time.
The client can abort a response while the transaction is still in progress if it
decides that the remainder of the data is not of immediate interest. In HTTP,
aborting the transaction requires closing the connection, which interferes with
the strategy of caching the open connection between transactions.
The client can request retrieval of portions of a document by specifying one or
more byte ranges (by offset and count) in the HTTP request headers. Each
range can be relative to either the beginning or the end of the file. The client
can specify as many ranges as it wants in the request, and the response consists
of multiple blocks, each properly tagged.
The client can initiate multiple concurrent transactions in an attempt to ob-
tain multiple responses in parallel. This is commonly done, for instance, to re-
trieve inline images referenced from an HTML document. This strategy is not
1024
APPENDIX F
Linearized PDF
always reliable and may backfire if the transactions interfere with each other
by competing for scarce resources in the server or the communication chan-
nel.
Note: Extensive experimentation has determined that having multiple concurrent
transactions does not work very well for PDF in some important environments.
Therefore, Linearized PDF is designed to enable good performance to be achieved
using only one transaction at a time. In particular, this means that the client must
have sufficient information to determine the byte ranges for all the objects re-
quired to display a given page of the PDF file so that it can specify all those byte
ranges in a single request.
The following additional assumptions are made about the PDF viewer application
and its local environment:
The viewer application has plenty of local temporary storage available. It
should rarely need to retrieve a given portion of a PDF document more than
once from the server.
The viewer application is able to display PDF data quickly once it has been
received. The performance bottleneck is assumed to be in the transport sys-
tem (throughput or round-trip delay), not in the processing of data after it ar-
rives.
The consequence of these assumptions is that it may be advantageous for the cli-
ent to do considerable extra work to minimize delays due to communications.
Such work includes maintaining local caches and reordering actions according to
when the needed data becomes available.
F.2
Linearized PDF Document Structure
Except as noted below, all elements of a Linearized PDF file are as specified in
Section 3.4, “File Structure,” and all indirect objects in the file are numbered
sequentially in two groups, based on their order of appearance in the file.
The first group consists of the document catalog, certain other document-level
objects, and all objects belonging to the first page of the document. These ob-
jects are numbered sequentially, starting at the first object number after the last
number of the second group. (The stream containing the hint tables, called a
1025
SECTION F.2
Linearized PDF Document Structure
hint stream, may be numbered out of sequence; see Section F.2.5, “Hint Streams
(Parts 5 and 10).”)
The second group consists of all remaining objects in the document, including
all pages after the first, all shared objects (objects referenced from more than
one page, not counting objects referenced from the first page), and so forth.
These objects are numbered sequentially starting at 1.
These groups of objects are indexed by exactly two cross-reference table sections,
located as shown in Example F.1. The composition of these groups is discussed in
more detail in the sections that follow (ordered by the part number as shown in
this example, with one section for parts 5 and 10). All objects have a generation
number of 0.
Beginning with PDF 1.5, PDF files may contain object streams (see Section 3.4.6,
“Object Streams”). In linearized files containing object streams, the following
conditions apply:
Certain additional objects cannot be contained in an object stream: the linear-
ization dictionary, the document catalog, and page objects.
Objects stored within object streams are given the highest range of object num-
bers within the main and first-page cross-reference sections.
For files containing object streams, hint data can specify the location and size
of the object streams only (or uncompressed objects), not the individual com-
pressed objects. Similarly, shared object references should be made to the ob-
ject stream containing a compressed object, not to the compressed object itself.
Cross-reference streams (Section 3.4.7, “Cross-Reference Streams”) can be
used in place of traditional cross-reference tables. The logic described in this
chapter still applies, with the appropriate syntactic changes.
Example F.1
Part 1: Header
% PDF−1 . 1
% … Binary characters
1026
APPENDIX F
Linearized PDF
Part 2: Linearization parameter dictionary
43 0 obj
<< /Linearized 1.0
% Version
/L
54567
% File length
/H [ 475 598 ]
% Primary hint stream offset and length (part 5)
/O 45
% Object number of first page’s page object (part 6)
/E 5437
% Offset of end of first page
/N 11
% Number of pages in document
/T 52786
% Offset of first entry in main cross-reference table (part 11)
>>
endobj
Part 3: First-page cross-reference table and trailer
xref
43 14
0000000052 00000 n
0000000392 00000 n
0000001073 00000 n
Cross-reference entries for remaining objects in the first page
0000000475 00000 n
trailer
<< /Size 57
% Total number of cross-reference table entries in document
/Prev 52776
% Offset of main cross-reference table (part 11)
/Root 44 0 R
% Indirect reference to catalog (part 4)
Any other entries, such as Info and Encrypt
% (part 9)
>>
startxref
0
% Dummy cross-reference table offset
% % EOF
Part 4: Document catalog and other required document-level objects
44 0 obj
<< /Type /Catalog
/Pages 42 0 R
>>
endobj
… Other objects…
1027
SECTION F.2
Linearized PDF Document Structure
Part 5: Primary hint stream (may precede or follow part 6)
56 0 obj
<< /Length 457
Possibly other stream attributes, such as Filter
/S 221
% Position of shared object hint table
Possibly entries for other hint tables
>>
stream
Page offset hint table
Shared object hint table
Possibly other hint tables
endstream
endobj
Part 6: First-page section (may precede or follow part 5)
45 0 obj
<< /Type /Page
>>
endobj
… Outline hierarchy (if the PageMode value in the document catalog is UseOutlines)…
… Objects for first page, including both shared and nonshared objects…
Part 7: Remaining pages
1 0 obj
<< /Type /Page
Other page attributes, such as MediaBox, Parent, and Contents
>>
endobj
… Nonshared objects for this page…
… Each successive page followed by its nonshared objects…
… Last page followed by its nonshared objects…
Part 8: Shared objects for all pages except the first
… Shared objects…
Part 9: Objects not associated with pages, if any
… Other objects…
1028
APPENDIX F
Linearized PDF
Part 10: Overflow hint stream (optional)
… Overflow hint stream…
Part 11: Main cross-reference table and trailer
xref
0 43
0000000000 65535 f
Cross-reference entries for all except first page’s objects
trailer
<< /Size 43 >>
% Trailer need not contain other entries; in particular,
startxref
% it should not have a Prev entry
257
% Offset of first-page cross-reference table (part 3)
% % EOF
F.2.1
Header (Part 1)
The Linearized PDF file begins with the standard header line (see Section 3.4.1,
“File Header”). Linearization is independent of PDF version number and can be
applied to any PDF file of version 1.1 or greater.
The binary characters following the percent sign on the second line are characters
with codes 128 or greater, as recommended in Section 3.4.1, “File Header.”
F.2.2
Linearization Parameter Dictionary (Part 2)
Following the header, the first object in the body of the file (part 2) must be an in-
direct dictionary object, the linearization parameter dictionary, containing the pa-
rameters listed in Table F.1. All values in this dictionary must be direct objects.
There are no references to this dictionary anywhere in the document; however,
the first-page cross-reference table (Part 3) contains a normal entry for it.
The linearization parameter dictionary must be entirely contained within the first
1024 bytes of the PDF file. This limits the amount of data a viewer application
must read before deciding whether the file is linearized.
1029
SECTION F.2
Linearized PDF Document Structure
TABLE F.1 Entries in the linearization parameter dictionary
PARAMETER
TYPE
VALUE
Linearized
number
(Required) A version identification for the linearized format. As usual, a
change in the integer part indicates an incompatible change in the linearized
format, while a change in the fractional part indicates a backward-compatible
change. The current version is 1.0.
L
integer
(Required) The length of the entire file in bytes. It must be exactly equal to the
actual length of the PDF file. A mismatch indicates that the file is not
linearized and must be treated as ordinary PDF, ignoring linearization in-
formation. (If the mismatch resulted from appending an update, the linear-
ization information may still be correct but requires validation; see Section
F.4.6, “Accessing an Updated File,” for details.)
H
array
(Required) An array of two or four integers,
[ offset1 length1 ] or
[ offset1 length1 offset2 length2 ]. offset1 is the offset of the primary hint
stream from the beginning of the file. (This is the beginning of the stream ob-
ject, not the beginning of the stream data.) length1 is the length of this stream,
including stream object overhead.
If the value of the primary hint stream dictionary’s Length entry is an indirect
reference, the object it refers to must immediately follow the stream object,
and length1 also includes the length of the indirect length object, including
object overhead. (See implementation note 178 in Appendix H.)
If there is an overflow hint stream, offset2 and length2 specify its offset and
length. (See implementation note 179 in Appendix H.)
O
integer
(Required) The object number of the first page’s page object.
E
integer
(Required) The offset of the end of the first page (the end of part 6 in Example
F.1), relative to the beginning of the file. (See implementation note 180 in
Appendix H.)
N
integer
(Required) The number of pages in the document.
1030
APPENDIX F
Linearized PDF
PARAMETER
TYPE
VALUE
T
integer
(Required) In documents that use standard main cross-reference tables (in-
cluding hybrid-reference files; see “Compatibility with Applications That Do
Not Support PDF 1.5” on page 109), this entry represents the offset of the
white-space character preceding the first entry of the main cross-reference
table (the entry for object number 0), relative to the beginning of the file.
Note that this differs from the Prev entry in the first-page trailer, which gives
the location of the xref line that precedes the table.
In PDF 1.5 and later documents that use cross-reference streams exclusively
(see Section 3.4.7, “Cross-Reference Streams”), this entry represents the off-
set of the main cross-reference stream object.
P
integer
(Optional) The page number of the first page (see Section F.2.6, “First-Page
Section (Part 6)”). Default value: 0.
F.2.3
First-Page Cross-Reference Table and Trailer (Part 3)
Part 3 contains the cross-reference table for objects belonging to the first page
(discussed in Section F.2.6, “First-Page Section (Part 6)”) as well as for the docu-
ment catalog and document-level objects appearing before the first page (dis-
cussed in Section F.2.4, “Document Catalog and Document-Level Objects (Part
4)”). Additionally, this cross-reference table contains entries for the linearization
parameter dictionary (at the beginning) and the primary hint stream (at the end).
This table is a valid cross-reference table as defined in Section 3.4.3, “Cross-
Reference Table,” although its position in the file is unconventional. It consists of
a single cross-reference subsection that has no free entries.
Note: In PDF 1.5 and later, cross-reference streams (see Section 3.4.7, “Cross-Refer-
ence Streams”) may be used in linearized files in place of traditional cross-reference
tables. The logic described in this section, along with the appropriate syntactic
changes for cross-reference streams, still applies.
Below the table is the first-page trailer. The trailer’s Prev entry gives the offset of
the main cross-reference table near the end of the file. This is valid PDF syntax,
although the trailers are linked in an unusual order. A PDF viewer application
that is unaware of linearization interprets the first-page cross-reference table as
an update to an original document that is indexed by the main cross-reference ta-
ble.
1031
SECTION F.2
Linearized PDF Document Structure
The first-page trailer must contain valid Size and Root entries, as well as any
other entries needed to display the document. The Size value must be the com-
bined number of entries in both the first-page cross-reference table and the main
cross-reference table.
The first-page trailer may optionally end with startxref, an integer, and %%EOF,
just as in an ordinary trailer. This information is ignored.
F.2.4
Document Catalog and Document-Level Objects (Part 4)
Following the first-page cross-reference table and trailer are the catalog dictio-
nary and other objects that are required when the document is opened. These
additional objects (constituting part 4) include the values of the following entries
if they are present and are indirect objects:
The ViewerPreferences entry in the catalog.
The PageMode entry in the catalog. Note that if the value of PageMode is
UseOutlines, the outline hierarchy is located in part 6; otherwise, the outline
hierarchy, if any, is located in part 9. See Section F.2.9, “Other Objects (Part 9)”
for details.
The Threads entry in the catalog, along with all thread dictionaries it refers to.
This does not include the threads’ information dictionaries or the individual
bead dictionaries belonging to the threads.
The OpenAction entry in the catalog.
The AcroForm entry in the catalog. Only the top-level interactive form dictio-
nary is needed, not the objects that it refers to.
The Encrypt entry in the first-page trailer dictionary. All values in the encryp-
tion dictionary must also be located here.
Objects that are not ordinarily needed when the document is opened should not
be located here but instead should be at the end of the file; see Section F.2.9, “Oth-
er Objects (Part 9).” This includes objects such as page tree nodes, the document
information dictionary, and the definitions for named destinations.
Note that the objects located here are indexed by the first-page cross-reference
table, even though they are not logically part of the first page.
1032
APPENDIX F
Linearized PDF
F.2.5
Hint Streams (Parts 5 and 10)
The core of the linearization information is stored in data structures known as
hint tables, whose format is described in Section F.3, “Hint Tables.” They provide
indexing information that enables the client to construct a single request for all
the objects that are needed to display any page of the document or to retrieve cer-
tain other information efficiently. The hint tables may contain additional infor-
mation to optimize access by plug-in extensions to application-specific data.
The hint tables are not logically part of the information content of the document;
they can be derived from the document. Any action that changes the document—
for instance, appending an incremental update—invalidates the hint tables. The
document remains a valid PDF file but is no longer linearized; see Section F.4.6,
“Accessing an Updated File,” for details.
The hint tables are binary data structures that are enclosed in a stream object.
Syntactically, this stream is a normal PDF indirect object. However, there are no
references to the stream anywhere in the document. Therefore, it is not logically
part of the document, and an operation that regenerates the document may re-
move the stream.
Usually, all the hint tables are contained in a single stream, known as the primary
hint stream. Optionally, there may be an additional stream containing more hints,
known as the overflow hint stream. The contents of the two hint streams are to be
concatenated and treated as if they were a single unbroken stream.
The primary hint stream, which is required, is shown as part 5 in Example F.1.
The order of this part and the first-page section, shown as part 6, may be re-
versed; see Section F.4, “Access Strategies,” for considerations on the choice of
placement. The overflow hint stream, part 10, is optional. (See implementation
note 179 in Appendix H.)
The location and length of the primary hint stream, and of the overflow hint
stream if present, are given in the linearization parameter dictionary at the begin-
ning of the file.
The hint streams are assigned the last object numbers in the file—that is, after the
object number for the last object in the first page. Their cross-reference table
entries are at the end of the first-page cross-reference table. This object number
assignment is independent of the physical locations of the hint streams in the file.
1033
SECTION F.2
Linearized PDF Document Structure
(This convention keeps their object numbers from conflicting with the number-
ing of the linearized objects.)
With one exception, the values of all entries in the hint streams’ dictionaries must
be direct objects and can contain no indirect object references. The exception is
the stream dictionary’s Length entry (see the discussion of the H entry in Table
F.1).
In addition to the standard stream attributes, the dictionary of the primary hint
stream contains entries giving the position of the beginning of each hint table in
the stream. These positions are given in bytes relative to the beginning of the
stream data (after decoding filters, if any, are applied) and with the overflow
hint stream concatenated if present. The dictionary of the overflow hint stream
should not contain these entries. The keys designating the standard hint tables
in the primary hint stream’s dictionary are listed in Table F.2; Section F.3, “Hint
Tables,” documents the format of these hint tables. Additionally, there is a re-
quired page offset hint table, which must be the first table in the stream and
must start at offset 0.
TABLE F.2 Standard hint tables
KEY
HINT TABLE
S
(Required) Shared object hint table (see Section F.3.2, “Shared Object Hint Ta-
ble”)
T
(Present only if thumbnail images exist) Thumbnail hint table (see Section F.3.3,
“Thumbnail Hint Table”)
O
(Present only if a document outline exists) Outline hint table (see Section F.3.4,
“Generic Hint Tables”)
A
(Present only if article threads exist) Thread information hint table (see Section
F.3.4, “Generic Hint Tables”)
E
(Present only if named destinations exist) Named destination hint table (see
Section F.3.4, “Generic Hint Tables”)
V
(Present only if an interactive form dictionary exists) Interactive form hint table
(see Section F.3.5, “Extended Generic Hint Tables”)
I
(Present only if a document information dictionary exists) Information dictio-
nary hint table (see Section F.3.4, “Generic Hint Tables”)
1034
APPENDIX F
Linearized PDF
KEY
HINT TABLE
C
(Present only if a logical structure hierarchy exists; PDF 1.3) Logical structure
hint table (see Section F.3.5, “Extended Generic Hint Tables”)
L
(PDF 1.3) Page label hint table (see Section F.3.4, “Generic Hint Tables”)
R
(Present only if a renditions name tree exists; PDF 1.5) Renditions name tree
hint table (see Section F.3.5, “Extended Generic Hint Tables”)
B
(Present only if embedded file streams exist; PDF 1.5) Embedded file stream hint
table (see Section F.3.6, “Embedded File Stream Hint Tables”)
New keys may be registered for additional hint tables required for new PDF
features or for application-specific data accessed by plug-in extensions. See
Appendix E for further information.
F.2.6
First-Page Section (Part 6)
As mentioned earlier, the section containing objects belonging to the first page of
the document may either precede or follow the primary hint stream. The starting
file offset and length of this section can be determined from the hint tables. In
addition, the E entry in the linearization parameter dictionary specifies the end of
the first page (as an offset relative to the beginning of the file), and the O entry
gives the object number of the first page’s page object.
This part of the file contains all the objects needed to display the first page of the
document. Ordinarily, the first page is page 0—that is, the leftmost leaf page node
in the page tree. However, if the document catalog contains an OpenAction entry
that specifies opening at some page other than page 0, that page is considered the
first page and should be located here. The page number of the first page is given
in the P entry of the linearization parameter dictionary. (See also implementation
note 181 in Appendix H.)
The following objects should be contained in the first-page section:
The page object for the first page. This object must be the first one in this part
of the file. Its object number is given in the linearization parameter dictionary.
This page object must explicitly specify all required attributes, such as
1035
SECTION F.2
Linearized PDF Document Structure
Resources and MediaBox; the attributes cannot be inherited from ancestor page
tree nodes.
The entire outline hierarchy, if the value of the PageMode entry in the catalog
is UseOutlines. (If the PageMode entry is omitted or has some other value and
the document has an outline hierarchy, the outline hierarchy appears in part 9;
see Section F.2.9, “Other Objects (Part 9)” for details.)
All objects that the page object refers to, to an arbitrary depth, except page tree
nodes or other page objects. This includes objects referred to by its Contents,
Resources, Annots, and B entries, but not the Thumb entry.
The order of objects referenced from the page object should facilitate early user
interaction and incremental display of the page data as it arrives. The following
order is recommended:
1.
The Annots array and all annotation dictionaries, to a depth sufficient for
those annotations to be activated. Information required to draw the annota-
tion can be deferred until later since annotations are always drawn on top of
(hence after) the contents.
2.
The B (beads) array and all bead dictionaries, if any, for this page. If any beads
exist for this page, the B array is required to be present in the page dictionary.
Additionally, each bead in the thread (not just the first bead) must contain a T
entry referring to the associated thread dictionary.
3.
The resource dictionary, but not the resource objects contained in the dic-
tionary.
4.
Resource objects, other than the types listed below, in the order that they are
first referenced (directly or indirectly) from the content stream. If the contents
are represented as an array of streams, each resource object should precede the
stream in which it is first referenced. Note that Font, FontDescriptor, and
Encoding resources should be included here, but not substitutable font files
referenced from font descriptors (see item 7 below).
5.
The page contents (Contents). If large, this should be represented as an array
of indirect references to content streams, which in turn are interleaved with
the resources they require. If small, the entire contents should be a single con-
tent stream preceding the resources.
6.
Image XObjects, in the order that they are first referenced. Images are assumed
to be large and slow to transfer; therefore, the viewer application defers ren-
dering images until all the other contents have been displayed.
1036
APPENDIX F
Linearized PDF
7. FontFile streams, which contain the actual definitions of embedded fonts.
These are assumed to be large and slow to transfer; therefore, the viewer appli-
cation uses substitute fonts until the real ones have arrived. Only those fonts
for which substitution is possible can be deferred in this way. (Currently, this
includes any Type 1 or TrueType font that has a font descriptor with the
Nonsymbolic flag set, indicating the Adobe standard Latin character set).
See Section F.4, “Access Strategies,” for additional discussion about object order
and incremental drawing strategies.
F.2.7
Remaining Pages (Part 7)
Part 7 of the Linearized PDF file contains the page objects and nonshared objects
for all remaining pages of the file, with the objects for each page grouped togeth-
er. The pages are contiguous and are ordered by page number. If the first page of
the file is not page 0, this section starts with page 0 and skips over the first page
when its position in the sequence is reached.
For each page, the objects required to display that page are grouped together,
except for resources and other objects that are shared with other pages. Shared
objects are located in the shared objects section (part 8). The starting file offset
and length of any page can be determined from the hint tables.
The recommended order of objects within a page is essentially the same as in the
first page. In particular, the page object must be the first object in each section.
In most cases, unlike for the first page, little benefit is gained from interleaving
contents with resources because most resources other than images—fonts in par-
ticular—are shared among multiple pages and therefore reside in the shared ob-
jects section. Image XObjects usually are not shared, but they should appear at
the end of the page’s section of the file, since rendering of images is deferred.
F.2.8
Shared Objects (Part 8)
Part 8 of the file contains objects, primarily named resources, that are referenced
from more than one page but that are not referenced (directly or indirectly) from
the first page. The hint tables contain an index of these objects. For more infor-
mation on named resources, see Section 3.7.2, “Resource Dictionaries.”
1037
SECTION F.2
Linearized PDF Document Structure
The order of these objects is essentially arbitrary. However, wherever a resource
consists of a multiple-level structure, all components of the structure should be
grouped together. If only the top-level object is referenced from outside the
group, the entire group can be described by a single entry in the shared object
hint table. This helps to minimize the size of the shared object hint table and the
number of individual references from entries in the page offset hint table. (See
also implementation note 182 in Appendix H.)
F.2.9
Other Objects (Part 9)
Following the shared objects are any other objects that are part of the document
but are not required for displaying pages. These objects are divided into function-
al categories. Objects within each of these categories should be grouped together;
the relative order of the categories is unimportant.
The page tree. This object can be located in this section because the viewer ap-
plication never needs to consult it. Note that all Resources attributes and other
inheritable attributes of the page objects must be pushed down and replicated
in each of the leaf page objects (but they may contain indirect references to
shared objects).
Thumbnail images. These objects should simply be ordered by page number.
(The thumbnail image for page 0 should be first, even if the first page of the
document is some page other than 0.) Each thumbnail image consists of one or
more objects, which may refer to objects in the thumbnail shared objects sec-
tion (see the next item).
Thumbnail shared objects. These are objects that are shared among some or all
thumbnail images and are not referenced from any other objects.
The outline hierarchy, if not located in part 6. The order of objects should be the
same as the order in which they are displayed by the viewer application. This is
a preorder traversal of the outline tree, skipping over any subtree that is closed
(that is, whose parent’s Count value is negative). Following that should be the
subtrees that were skipped over, in the order in which they would have ap-
peared if they were all open.
Thread information dictionaries, referenced from the I entries of thread dictio-
naries. Note that the thread dictionaries themselves are located with the docu-
ment catalog and the bead dictionaries with the individual pages.
1038
APPENDIX F
Linearized PDF
Named destinations. These objects include the value of the Dests or Names
entry in the document catalog and all the destination objects that it refers to.
See Section F.4.2, “Opening at an Arbitrary Page.”
The document information dictionary and the objects contained within it.
The interactive form field hierarchy. This group of objects does not include the
top-level interactive form dictionary, which is located with the document cata-
log.
Other entries in the document catalog that are not referenced from any page.
(PDF 1.3) The logical structure hierarchy.
(PDF 1.5) The renditions name tree hierarchy.
(PDF 1.5) Embedded file streams.
F.2.10
Main Cross-Reference and Trailer (Part 11)
Part 11 is the cross-reference table for all objects in the PDF file except those
listed in the first-page cross-reference table (part 3). As indicated earlier, this
cross-reference table plays the role of the original cross-reference table for the file
(before any updates are appended) and must conform to the following rules:
It consists of a single cross-reference subsection, beginning at object number 0.
The first entry (for object number 0) must be a free entry.
The remaining entries are for in-use objects, which are numbered consecutive-
ly, starting at 1.
The startxref line gives the offset of the first-page cross-reference table. The Prev
entry of the first-page trailer gives the offset of the main cross-reference table.
The main trailer has no Prev entry and in fact does not need to contain any en-
tries other than Size.
Note: In PDF 1.5 and later, cross-reference streams (see Section 3.4.7, “Cross-Refer-
ence Streams”) may be used in linearized files in place of traditional cross-reference
tables. The logic described in this chapter, along with the appropriate syntactic
changes for cross-reference streams, still applies.
1039
SECTION F.3
Hint Tables
F.3
Hint Tables
The core of the linearization information is stored in two or more hint tables, as
indicated by the attributes of the primary hint stream (see Section F.2.5, “Hint
Streams (Parts 5 and 10)”). The format of the standard hint tables is described in
this section.
There can be additional hint tables for application-specific data that is accessed
by plug-in extensions. A generic format for such hint tables is defined; see Section
F.3.4, “Generic Hint Tables.” Alternatively, the format of a hint table can be private
to the application; see Appendix E for further information.
Each hint table consists of a portion of the stream, beginning at the position in
the stream indicated by the corresponding stream attribute. Additionally, there is
a required page offset hint table, which must be the first table in the stream and
must start at offset 0. (If there is an overflow hint stream, its contents are to be ap-
pended seamlessly to the primary hint stream; hint table positions are relative to
the beginning of this combined stream.) In general, this byte stream is treated as a
bit stream, high-order bit first, which is then subdivided into fields of arbitrary
width without regard to byte boundaries. However, each hint table begins at a
byte boundary.
The hint tables are designed to encode the required information as compactly as
possible. Interpreting the hint tables requires reading them sequentially; they are
not designed for random access. The client is expected to read and decode the
tables once and retain the information for as long as the document remains open.
A hint table encodes the positions of various objects in the file. The representa-
tion is either explicit (an offset from the beginning of the file) or implicit (accu-
mulated lengths of preceding objects). Regardless of the representation, the
resulting positions must be interpreted as if the primary hint stream itself were
not present. That is, a position greater than the hint stream offset must have the
hint stream length added to it to determine the actual offset relative to the begin-
ning of the file. (The hint stream offset and hint stream length are the values
offset1 and length1 in the H array in the linearization parameter dictionary at the
beginning of the file.)
The reason for this rule is that the length of the primary hint stream depends on
the information contained within the hint tables, which is not known until after
1040
APPENDIX F
Linearized PDF
they have been generated. Any information contained in the hint tables must not
depend on knowing the primary hint stream’s length in advance.
Note that this rule applies only to offsets given in the hint tables and not to offsets
given in the cross-reference tables or linearization parameter dictionary. Also, the
offset and length of the overflow hint stream, if present, need not be taken into
account, since this object follows all other objects in the file.
Note: In linearized files that use object streams (Section 3.4.6, “Object Streams), the
position specified in a hint table for a compressed object is to be interpreted as a byte
range in which the object can be found, not as a precise offset. Viewer applications
should locate the object via a cross-reference stream, as it would if the hint table
were not present.
F.3.1
Page Offset Hint Table
The page offset hint table provides information required for locating each page.
Additionally, for each page except the first, it also enumerates all shared objects
that the page references, directly or indirectly.
This table begins with a header section, described in Table F.3, followed by one or
more per-page entries, described in Table F.4. Note that the items making up each
per-page entry are not contiguous; they are broken up with items from entries for
other pages. The order of items making up the per-page entries is as follows:
1. Item 1 for all pages, in page order starting with the first page
2. Item 2 for all pages, in page order starting with the first page
3. Item 3 for all pages, in page order starting with the first page
4. Item 4 for all shared objects in the second page, followed by item 4 for all
shared objects in the third page, and so on
5. Item 5 for all shared objects in the second page, followed by item 5 for all
shared objects in the third page, and so on
6. Item 6 for all pages, in page order starting with the first page
7. Item 7 for all pages, in page order starting with the first page
1041
SECTION F.3
Hint Tables
Note: All the items in Table F.3 that specify a number of bits needed, such as item 3,
can have values in the range 0 through 32. Although that range requires only 6 bits,
16-bit numbers are used.
TABLE F.3 Page offset hint table, header section
ITEM
SIZE (BITS)
DESCRIPTION
1
32
The least number of objects in a page (including the page object itself).
2
32
The location of the first page’s page object.
3
16
The number of bits needed to represent the difference between the greatest
and least number of objects in a page.
4
32
The least length of a page in bytes. This is the least length from the beginning
of a page object to the last byte of the last object used by that page.
5
16
The number of bits needed to represent the difference between the greatest
and least length of a page, in bytes.
6
32
The least offset of the start of any content stream, relative to the beginning of
its page. (See implementation note 183 in Appendix H.)
7
16
The number of bits needed to represent the difference between the greatest
and least offset to the start of the content stream. (See implementation note
183 in Appendix H.)
8
32
The least content stream length. (See implementation note 184 in Appendix
H.)
9
16
The number of bits needed to represent the difference between the greatest
and least content stream length. (See implementation note 184 in Appendix
H.)
10
16
The number of bits needed to represent the greatest number of shared object
references.
11
16
The number of bits needed to represent the numerically greatest shared ob-
ject identifier used by the pages (discussed further in Table F.4, item 4).
1042
APPENDIX F
Linearized PDF
ITEM
SIZE (BITS)
DESCRIPTION
12
16
The number of bits needed to represent the numerator of the fractional posi-
tion for each shared object reference. For each shared object referenced from
a page, there is an indication of where in the page’s content stream the object
is first referenced. That position is given as the numerator of a fraction,
whose denominator is specified once for the entire document (in the next
item in this table). The fraction is explained in more detail in Table F.4,
item 5.
13
16
The denominator of the fractional position for each shared object reference.
TABLE F.4 Page offset hint table, per-page entry
ITEM
SIZE (BITS)
DESCRIPTION
1
See Table F.3, item 3
A number that, when added to the least number of objects in a page (Table
F.3, item 1), gives the number of objects in the page. The first object of the
first page has an object number that is the value of the O entry in the
linearization parameter dictionary at the beginning of the file. The first
object of the second page has an object number of 1. Object numbers for sub-
sequent pages can be determined by accumulating the number of objects in
all previous pages.
2
See Table F.3, item 5
A number that, when added to the least page length (Table F.3, item 4), gives
the length of the page in bytes. The location of the first object of the first page
can be determined from its object number (the O entry in the linearization
parameter dictionary) and the cross-reference table entry for that object (see
Section F.2.3, “First-Page Cross-Reference Table and Trailer (Part 3)”). The
locations of subsequent pages can be determined by accumulating the lengths
of all previous pages. Note that it is necessary to skip over the primary hint
stream, wherever it is located.
3
See Table F.3, item 10
The number of shared objects referenced from the page. For the first page,
this number must be 0; the next two items start with the second page.
4
See Table F.3, item 11
(One item for each shared object referenced from the page) A shared object
identifier—that is, an index into the shared object hint table (described in
Section F.3.2, “Shared Object Hint Table”). Note that a single entry in the
shared object hint table can designate a group of shared objects, only one of
which is referenced from outside the group. That is, shared object identifiers
are not directly related to object numbers.
This identifier combines with the numerators provided in item 5 to form a
shared object reference.

Была ли эта страница вам полезна?
Да!Нет
5 посетителей считают эту страницу полезной.
Большое спасибо!
Ваше мнение очень важно для нас.

Нет комментариевНе стесняйтесь поделиться с нами вашим ценным мнением.

Текст

Политика конфиденциальности