Here is the Obama PDF hierarchy
- Document
- Catalog
- Pages
- Page 1
- Raster Image ProcSet
- Image 1
- Image 2
- Image 3
- Image 4
- Image 5
- Image 6
- Image 7
- Image 8
- Image 9
Here is the equivalent Illustrator document hierarchy after the Obama PDF is imported
- Document
- Layer 1
- Clip Group 0
- Clip Path
- Clip Group 1
- Clip Group 2
- Clip Group 3
- Clip Group 4
- Clip Group 5
- Clip Group 6
- Clip Group 7
- Clip Group 8
- Clip Group 9
By comparison, here's the structure of the optimized PDF of the
Photoshop User cover I scanned this morning.
- Document
- Catalog
- Pages
- Page 1
- Raster Image ProcSet
- Image 1
- Image 2
- Image 3
- Image 4
And here is the object hierarchy reported by Illustrator when I imported that PDF
- Document
- Layer 1
- Clip Group 0
- Clip Path
- Clip Group 1
- Clip Group 2
- Clip Group 3
- Clip Group 4
Now the reader might legitimately ask two questions.
First, why does Illustrator put each subimage into its own "Clip Group?"
The simple answer is because Illustrator is stupid. As Abbadon has lamented, importing PDFs into Illustrator is a losing proposition. "Clipping" in general computer graphics terms means displaying only part of a raster image. One "clips" the image to the viewport, for example, not displaying the parts of the picture that are not within the scrollable viewport in a window. One can also clip an image to an arbitrary shape, in cookie-cutter fashion.
In PDF terms, clipping occurs first by means of a rectangular viewport representing the printable and visible portions of the media, then optionally by means of a bitmask -- an image composed only of black-or-white pixels that determine which source pixels find their way into the destination raster. Because PDF drawing commands can reset the graphics state easily in the middle of rendering an object, there is a shortcut you can use in PDF to "mask out" the cumulative image. This is what was done in the Obama PDF, and what is commonly done as part of PDF optimization -- it replaces unnecessarily "deep" or unacceptably "noisy" portions of a scanned image with single-bit values that invoke the known value of the otherwise ephemeral graphics state.
In Illustrator terms, clipping is done at a higher geometric level. Underlying objects are clipped to the
outline of high-level objects (e.g., paths and text) according to standard set-operation combinations. Parts of objects that lie within the boundaries of other objects are said to be clipped by them. Illustrator puts these in a special type of group called a Clip Group because the computations that resolve the visible portions of an object among possibly dozens of object interactions are complex and time-consuming. Identifying that group for special handling allows Illustrator to cache the interaction results and store (behind the scenes) intermediate clip regions.
Illustrator
can use opacity masks to clip, but even these are typically created and stored as parameterized objects (e.g., gradient descriptors). You have to go to extreme lengths to get Illustrator to recognize an imported bitmap as an opacity map, and still the best method is for Illustrator to trace the bitmap as path object. Illustrator's importer defers that decision. But still when it sees a PDF object described as
/XObject /Subtype /Image /ImageMask true /BitsPerComponent 1, it recognizes it as something that eventually will need to be clipped in some fashion or another. Hence it creates a clipping group for it provisionally, even if the group ends up being degenerate in the Illustrator data model.
As elaborated in the PDF documentation, this is not a new Layer. This is Illustrator's best stab at representing a PDF object in a way that will allow it to be rendered and represented appropriately in Illustrator's dissimilar data structure and rendering model. It turns out the Clip Group contains only the clipping mask (to be interpreted either as a path-generator or an opacity mask) and simply modifies the accumulated group graphic state as the rendering pipeline proceeds.
Second, Robert's quote indeed mentions layered content; what does that hierarchy look like?
I have a simple but valid Illustrator document as follows
- Document
- Registrations Layer
- Crop Mark Group
- Crop Mark 1
- Crop Mark 2
- Crop Mark 3
- Crop Mark 4
- Registration Mark Group
- Registration Mark Cyan
- Registration Mark Yellow
- Registration Mark Blue
- Registration Mark Black
- Raster Content Layer
- Image 1 Group
- Production Still Photo
- Text Object for Caption 1
- Image 2 Group
- Cast Thumbnails
- Text Object for Caption 2
- Vector Content Layer
- Rectangle Gradient Object
- Text Object for Cast List and Production Credits
- Title Group
- Text Object for DVD Title
- Horizontal Rule 1
- Horizontal Rule 2
In the Illustrator Layers pane, each of those entities will have its own line item entry, but only three of them are Layers. The rest are either user-defined groups or displayable objects I've created within those layers and groups.
Exporting to PDF, I get this:
- Document
- Catalog
- Pages
- Page 1
- Raster Image ProcSet, CYMK color space, 8 bits per component
- Production Still Photo
- Cast Thumbnails
- Text ProcSet, Font Showtime
- Caption Text 1
- Caption Text 2
- Cast List and Production Credits Text
- Text ProcSet, Font LaFleur
- PDF Drawing ProcSet
- Crop Mark 1
- Crop Mark 2
- Crop Mark 3
- Crop Mark 4
- Registration Mark Cyan
- Registration Mark Yellow
- Registration Mark Blue
- Registration Mark Black
- Rectangle Gradient Object
- Text Box
- Horizontal Rule 1
- Horisontal Rule 2
Now as a tangential exercise, here's what I get when I import that PDF back into Illustrator
- Document
- Layer 1
- Clip Group 0
- Clipping Path
- Crop Mark 1
- Crop Mark 2
- Crop Mark 3
- Crop Mark 4
- Registration Mark Cyan
- Registration Mark Yellow
- Registration Mark Blue
- Registration Mark Black
- Production Still Photo
- Text Object for Caption 1
- Cast Thumbnails
- Text Object for Caption 2
- Rectangle Gradient Object
- Text Object for Cast List and Production Credits
- Text Object for DVD Title
- Horizontal Rule 1
- Horizontal Rule 2
Why aren't my layers and groups preserved? Because PDF doesn't know how to represent them effectively. So everything gets imported into a single Illustrator layer.
Now if I open the PDF in Adobe Acrobat Professional 9, I can select a Layers pane on the left. It initially shows me zero layers. That is, there is no "default layer" in PDF land. I can create the first layer ("Jay's New Layer"), but to add content to it I can only import it in the form of another PDF document. So if I do that by importing a simple multimedia document (the data sheet for an accelerometer) into my newly-created Layer in Acrobat.
- Document
- Catalog
- Pages
- Page 1
- Raster Image ProcSet
- PDF Drawing ProcSet
- Product Diagram
- Product Table Outline
- Text ProcSet, font Garamond
Saving the result as a new PDF, I get this:
- Document
- Catalog
- Pages
- Page 1
- Raster Image ProcSet, CYMK color space, 8 bits per component
- Production Still Photo
- Cast Thumbnails
- Product Illustration
- Text ProcSet, Font Showtime
- Caption Text 1
- Caption Text 2
- Cast List and Production Credits Text
- Text ProcSet, Font LaFleur
- Text ProcSet
- PDF Drawing ProcSet
- Crop Mark 1
- Crop Mark 2
- Crop Mark 3
- Crop Mark 4
- Registration Mark Cyan
- Registration Mark Yellow
- Registration Mark Blue
- Registration Mark Black
- Rectangle Gradient Object
- Text Box
- Horizontal Rule 1
- Horisontal Rule 2
- Product Diagram
- Product Table Outline
- Optional Content Group - "Jay's New Layer", intent logical-grouping, visible, non-scalable
- External ref. to Product Illustration
- External ref. to Product Diagram
- External ref. to Product Table Outline
- (many external references to text objects)
The rendering order and grouping remain unchanged. We've only added a new section that says for some purpose we can treat those objects similarly. In this case I've made it so that they can be made visible or invisible as a group (i.e., in an interactive PDF viewer) and that they should not be scaled if the user zooms into the document (useful for annotations). The grouping reasons are literally infinite.
Import this "layered" PDF back into Illustrator. What do we get?
- Document
- Layer 1
- Clip Group 0
- Clipping Path
- Crop Mark 1
- Crop Mark 2
- Crop Mark 3
- Crop Mark 4
- Registration Mark Cyan
- Registration Mark Yellow
- Registration Mark Blue
- Registration Mark Black
- Production Still Photo
- Text Object for Caption 1
- Cast Thumbnails
- Text Object for Caption 2
- Rectangle Gradient Object
- Text Object for Cast List and Production Credits
- Text Object for DVD Title
- Horizontal Rule 1
- Horizontal Rule 2
- Product Illustration
- Product Diagram
- Product Table Outline
- (many text objects)
Why didn't Illustrator recognize the optional-content group as a layer? Because content-groups are not layers the way
Illustrator understands them.
Because an Illustrator document is ruthlessly hierarchical, an object can appear only in one group, series of enclosing groups, and layer. But in PDF land we can say this
- Content Group 1 "Jay's Layer"
- External ref. to Object 42
- External ref. to Object 79
- Content Group 2 "Robert's Layer"
- External ref. to Object 13
- External ref. to Object 79
It's perfectly valid PDF, but to try to import PDF content groups like this into Illustrator as layers or even ordinary object groups would require an object to be in more than one Layer. Which, naturally, Illustrator can't do. So programs that interpret PDF files generally ignore OCGs unless they can be absolutely sure they were the ones who created it in the first place.
The only program that will interpret generic OCGs is ... Acrobat. And why? Because Acrobat's data model is guaranteed to be the PDF data model, by design.
This dilemma exists because while OCGs
can represent original layers, they are not
required in PDF terms to do so. They are employed to serve whatever purpose the PDF author (not the original author) intends. Hence it is unwise for a PDF importer to assume that OCGs he encounters in PDF files will be compatible with his data model.
Oh, and this facility doesn't exist in PDF 1.3, the version that Obama's PDF is written in.
So let's revisit Robert's quote again. InDesign will let you create a number of printed
and interactive publications. If you're foolish enough to design a web site in InDesign, interactive content will be represented in a PDF export with OCGs having a certain intent setting. And that means when you go to print them, those stupid OCGs (which may derive from InDesign "Layers" that represented interaction panes) give you all the optional content at once -- which doesn't look good on paper.
That's why Robert's quote is limited to (a) PDFs exported from InDesign documents, (b) printing on paper, and (c) PDF versions that support optional content.