Question: Are all parts of the PDF file that contain image information called objects or is there a primary one that is called something else?
All parts of a PDF file are objects except for the cross-reference table ("xref"). The
xref simply locates the other objects within the file and can have no other function.
PDF objects are weakly typed and weakly composed. Weak typing means that an object can have an implicit type (e.g., an array of values or a single string value) determined by its contents, or may be declared explicitly as a supported PDF type or subtype. Weak composition means an object can have a single atomic value or may also include "dictionary" of name-value pairs that aggregates atomic types into a composite type. Objects may refer to other objects.
PDF document structure is implemented as a set of strongly composite types: a /Catalog object statically lists the pages the PDF file contains. For each page, a /Page object lists the displayable and non-displayable (e.g., font) objects that appear on the page. Those objects are often grouped into "procedure sets" (i.e., objects of type /ProcSet) that coalesce rendering parameters for displayable objects of similar type (e.g., to render all text objects using the same font, or all graphics objects using the same color model). The PDF imaging model (described in the aforementioned document) controls how these objects are interpreted for rendering.
Question: Does some PDF file creation software allow the user the ability to explicitly control where objects are generated or to build objects themselves and incorporate them into the PDF file?
Yes. However "PDF file creation software" is a broad topic. Within the scope of this investigation we consider PDF converters and PDF optimizers, which are the tools Birthers and their critics discuss, and the tools alleged to have been used in preparing the Obama PDF for publication.
PDF converters rewrite the logical content of some other data into equivalent PDF form. For example, Adobe Distiller reads Postscript data and produces a PDF that renders a visual equivalent to the PS. A better example is the PDF export features in all Adobe creative products and most other similar products from other sources. The PDF file format is fully specified and freely published. Hence any software developer has the tools available to translate his program's data into a PDF equivalent.
Adobe Photoshop primarily manipulates pixel maps. Those can be expressed as objects containing a byte stream and instructions for decoding it for the PDF imaging model. Adobe InDesign and Illustrator manipulate high-level descriptions of graphical primitives, such as a circle expressed in the form of the coordinate of its center and its radius, along with rendering hits such as fill properties and stroke properties. Those primitives can be expressed as equivalent PDF drawing primitives that are executed by the PDF renderer to produce output appropriate for the target device (e.g., the video screen, a laser printer, or a text-based terminal).
A number of parameters govern PDF conversion and exportation, derived from the myriad ways the desired visual result may be expressed in PDF terms. For example the exporter may render text as a pixel map using the authoring software's native fonts, rather than leave the rendering up to the eventual PDF renderer, because fine control can be achieved over the appearance of the text on all devices, where that matters. None of these parameters allows fine control over what objects are generated, especially those having to do with object aggregation and document structure. That said, whether the document includes option content (in the form of, say, hints about chapter headings) is parameterized in the export.
PDF optimizers rewrite a PDF stream in more optimal form, where "optimal" can have several different meanings within the scope of PDF rendering. In general it means simply "more desirable for some purpose such as speed of rendering, size of storage, fidelity of appearance, or compatibility in various environments."
For example, since PDF objects are versioned, a PDF optimizer may strip out deprecated object versions in order to shrink the file. In other cases the optimizer may create new objects that are simplified versions of existing objects, such as high-level graphics operators in place of pixel maps. Much has been said about the possibility of the typewritten text appearing as bitmapped glyphs as preparations for optical character-recognition. Whether that was accomplished, intended, or otherwise is largely irrelevant; for reasons laid out in the PDF reference, glyphs representing printable characters but encoded as pixel maps are optimized here as image masks to exploit a loophole in the PDF rendering model that makes it more compatible.
PDF optimizers have options to manipulate document structure complexity, which results in more or fewer objects being generated, or objects of semantically similar types to the original. But none of them gives especially fine control over the content of PDF objects.
In addition I have discussed special-purpose PDF editors, which I use to investigate the structure of PDFs and to manipulate their objects with fine control, such as to add annotations or change color models. These tools are the only ones capable of exercising the degree of control you allude to. For example I can change the name of the font called for in the file by editing an attribute displayed in a tree-type control. Examples of such tools include Adobe Acrobat Professional and the open-source
pdfedit. Neither Birthers nor their critics allege that any such tool was used to create or alter the Obama PDF.
PDF files can contain layers. The purpose of these layers is to store layers that were part of the source file for the image. The layers in the source file are the output of a graphics editing program.
No, this is not correct. There is no concept of "layers" in PDF that is an analogue of object groups and layers in Adobe-type creative products such as InDesign, Illustrator, Photoshop, and non-Adobe programs such as GIMP and AutoCAD. There is no type or subtype /Layer, nor any object type supported by PDF that mimics the behavior of a layer in those programs.
One might legitimately ask whether a /ProcSet could function as a layer. The answer is no. The /ProcSet object is a
procedural type, not an aggregation, which means it groups objects together based on the rendering methods that will be applied to it in the graphics pipeline, not whether they were aggregated in the authoring program (e.g., "Group Objects" in Illustrator) or included in the same logical layer.
The single /ProcSet in the Obama PDF contains a list of images to be rendered, all encoded in the same way and compressed with the same algorithm. The /ProcSet here simply provides a way for the same decoding and decompression algorithms to be applied in turn to a list of objects. If the page were to contain high-level text, another /ProcSet could group all the text objects together that were to be rendered using the same font.
The key to understanding this is to realize that objects wind up in a /ProcSet based on how the rendering pipeline software is meant to handle them, not what logical connection they may have held for the author. So a text /ProcSet would initially load the font, initialize the graphics context, then call the text-rendering algorithm on each text object referred to in the /ProcSet. Similarly a pixel-map /ProcSet initializes the color-space conversion, decompresses all the included pixel maps, decodes them into color-component values, and then creates device-specific output (e.g., writes RBG pixels into a frame buffer). And a graphics object for vector data also initializes the graphics state, then calls the vector-graphic rasterizer for each included set of drawing commands.
Hence consider a newsletter page in InDesign where one logical layer contains the cut and registration marks for a commercial print house, another logical layer contains the organization's logo in SVG form and SVG rules or other graphic design elements that never change from issue to issue, a third logical layer contains the text and pixel-map graphics that present the publication's masthead (also relatively unchanging), a fourth logical layer contains several large photographs and their captions, and the final layer contains body text imported from the organization's content-management system.
When exported to PDF, the exporter will not respect the logical layers ("logical" meaning that the grouping criteria means something to the author, such as separating static content from content that changes with each iteration). A /ProcSet will be generated to handle all the captions, presuming they are intended to appear in the same font. The associated photographs will not appear, even though they were created in the same logical layer and may have been "grouped" in InDesign for convenient dragging during composition. A different /ProcSet will handle all the body text, because it presumably uses a different font. The underlying semantic here is, "Treat all these objects the same way in the rendering pipeline." And a new /ProcSet can contain the PDF translation of all the SVG graphics, regardless of what original logical layer they appeared in. And finally all the pixel-map images can go into a /ProcSet with the same color model, encoding, etc.
But let's say the masthead bitmaps and the photograph bitmaps have a different color model. A PDF optimizer may convert one to the other's color model, for the simplicity of reducing the number of /ProcSets or for the fidelity of having fine software control over the color fidelity.
Similarly in AutoCAD I can create a "callout" (a label, in engineering terms, typically in a form governed by industry graphic standards). That callout object will be composed of a set of vector graphics (a box of a particular shape, with subdivisions, and a line connecting it to the called-out design feature) and one or more text objects. It is common to group those together so that within AutoCAD you can drag them around the screen as a group. But again, in PDF-land, the vector graphics go in one /ProcSet and the text goes into another.
Question: Are these layers used to render the image or are these layers stored in the PDF file purely to maintain the layer structure of the image for possible use in the future for graphic editing purposes?
The grouping of objects in PDF file is solely for the purposes of rendering. Adobe expressly warns PDF developers not to conflate the meaning of logical layers in their applications with rendering-based groupings that occur in PDF. Thus the groupings (I specifically avoid the term "layer" because Adobe specifically avoids it for this purpose) are purely to render the image and bear no relationship to any other logical layering.
Every /Page object, however, allows an optional /PieceInfo object to appear, either by reference or explicitly. The authoring software may place any information it wishes into a /PieceInfo entry. It has been suggested by Adobe that such a feature could store logical layer and grouping information that has meaning to the content-authoring software. It warns, however, that such information is ignored while rendering and has meaning only to the original authoring system, should it choose to try to interpret PDF data and reconstruct its original data types. This is apparently what InDesign uses, and what I was able to accomplish with Adobe Acrobat Professional.
So in the callout example above, the PDF document representing my engineering drawing could contain a /Page such as
Code:
5 o obj
<< /Type /Page
/Parent 3 0 R
/Resources [ 6 0 R 7 0 R ]
/Contents 4 0 R
/MediaBox [0 0 500 500]
/PieceInfo <<
/LastModified (D:20120909132500-0600)
/Private <<
/JayLayer1 <<
/Name (Callouts and Annotations)
/Objects [ 45 0 R 46 0 R 47 0 R ]
>>
/JayLayer2 <<
/Name (Electrical Cable Harnesses)
/Objects [ 110 0 R 111 0 R 124 0 R ]
>>
>>
>>
>>
endobj
6 0 obj
<< /ProcSet [ /PDF ] >>
<< /XObject << /CalloutBox 45 0 R /CalloutArrow 46 0 R /Dw01 110 0 R /Dw02 111 0 R 124 0 R>> >>
endobj
7 0 obj
<< /ProcSet [ /Text ] >>
<< /XObject << /CalloutText 47 0 R >> >>
endobj
In this example, everything within the scope of /Private is essentially unintelligible to anything except my program. I've determined that in this /Private section I will store the name of the layer I created in my program and a list of references to PDF objects that were derived from things I put in that layer. I have two layers in my original CAD drawing. One, titled "Callouts and Annotations" contains three objects: 45 (the rectangle around the text), 46 (the arrow connecting the box to the annotated feature), and 47 (the text inside the callout box). The other, titled "Electrical Cable Harneses," contains two objects representing electrical components in my design (objects 110, 110, and 124).
I can instruct my program to read this PDF file back in, and I can ask for the /Private data attached to the page, and I can interpret the data I've stored in my "dictionary" there, because I decided what it should contain. Adobe Acrobat, on the other hand, does not know what the values in my dictionary mean or how to interpret them. So it ignores them.
Object 5 is the /Page. It refers to two procedure resources, objects 6 and 7. These are /ProcSet objects. Object 6 identifies a procedure set of /PDF, indicating that I will be using PDF's graphical drawing primitives. The objects it refers to (and assigns symbolic names) are objects 45, 46, 110, 111, and 124, which are the line-drawing objects from all the original input data, regardless of layer.
Object 7 is another procedure set, requesting text output. The object is 47, the annotation inside the callout. Because it is a different
kind of data to be rendered, it is grouped in a different procedure set, even though it was logically contained in a layer with objects 45 and 46 in the original content-authoring CAD program, and may have been logically grouped together in that program for easy management. (I did not manufacture any grouping information for this example, only layering.)
Several articles on the web claimed that the Obama long form birth certificate PDF file contained multiple layers. Using standard PDF file terminology these articles used the wrong terminology. The Obama COB PDF file contained multiple objects. The Obama COB PDF file did not contain multiple layers.
Question: Is this correct?
Yes. The Birthers' pseudo-experts identified ordinary PDF objects (of type /XObject, subtype /Image) and referred to them as "layers." This is a fundamental error in understanding PDF data structures, rendering engines, and imaging models.
...it would be nice to find out how he voted on the first one.
I'm not seriously calling for another poll. The results of the previous poll should be sufficient to establish who has credibility and who does not.