Portable Document Format (.pdf)

ISO Standard

The Portable Document Format (PDF) is a ubiquitous digital paper standard designed to present documents consistently independent of application software, hardware, and operating systems.

Convertir PDF en DOCX

Convertisseur gratuit de PDF vers DOCX dans le navigateur. Convertissez des fichiers instantanément sur votre appareil.

Compresser des fichiers PDF en ligne

Compression 100% privée dans le navigateur – les fichiers ne quittent jamais votre appareil

Inspecter et métadonnées

Les systèmes d'exploitation et les analyseurs de fichiers identifient les fichiers PDF en inspectant la séquence d'octets binaires de tête :

Sélectionnez ou faites glisser des fichiers

Sélectionnez ou déposez des fichiers ici

Conversion 100 % privée dans le navigateur - les fichiers ne quittent jamais votre appareil

ou coller Ctrl+V
Garantie de confidentialité sans transmission réseau : 0 octet téléchargé vers des serveurs externes. Tout le traitement a lieu localement dans le bac à sable de votre navigateur.

Signature d'en-tête au niveau des octets (Magic Bytes)

Les systèmes d'exploitation et les analyseurs de fichiers identifient les fichiers PDF en inspectant la séquence d'octets binaires de tête :

SIGNATURE HEXADÉCIMALE (DÉCALAGE 0) :

25 50 44 46

REPRÉSENTATION ASCII : %PDF

Normalisation : ISO 32000-1 (2008) / ISO 32000-2 (2020)

Spécifications techniques

Architecture de conteneurBinary/ASCII hybrid object hierarchy with cross-reference table
CompressionFlateDecode (Deflate), DCTDecode (JPEG), CCITT Fax, JBIG2, JPXDecode
Boutisme des octetsBig-Endian (network byte order) for binary object streams
Espaces colorimétriquessRGB, Display P3, DeviceCMYK, Grayscale, Spot Colors (Pantone)
Canaux et structureMulti-layer vector, typography, raster streams, and embedded metadata
Dimensions maximales200 x 200 inches (UserUnit allows up to 15,000,000 inches in PDF 1.6+)
TransparenceFull alpha transparency, blend modes, and opacity masks
Streaming et progressifLinearized PDF ('Fast Web View') allows byte-range streaming via HTTP Range requests

Matrice de comparaison technique : PDF vs Concurrents

Attribut techniquePDF (Actuel)DOCXEPUBHTML5
Visual Layout FidelityPixel-perfect fixed geometric layout across all screens and printersFlow-based; varies across Word versions and installed fontsReflowable reader layout; no fixed typography guaranteesFluid CSS box model; screen-dependent rendering
Self-Containment100% self-contained with embedded fonts, raster streams, and vector assetsSelf-contained XML container (system fonts usually un-embedded)Self-contained ZIP container (packaged XHTML/CSS assets)Requires external HTTP resources unless packaged via MHTML
Cryptographic SignaturesNative PAdES and X.509 cryptographic digital signatures built into ISO specOffice XML DSig digital signatures supportedRequires custom digital container signaturesRequires Web Crypto API or external TLS transport guarantees
Mobile ResponsivenessFixed-page canvas (requires zooming and panning on small screens)Excellent responsive text reflowEngineered specifically for fluid mobile e-readingNative responsive layout with CSS media queries
StandardizationOpen International Standard (ISO 32000-2:2020)ECMA-376 / ISO/IEC 29500W3C / IDPF Open StandardW3C Living Standard / WHATWG

Modes de corruption courants et guide de récupération hexadécimale

⚠️ Acrobat or PDF viewer reports 'The file is damaged and could not be repaired' or opens completely blank.

Cause racine: FTP or HTTP transfer in ASCII mode converted binary byte sequences (replacing 0x0A with 0x0D 0x0A), mangling the cross-reference (xref) table byte offsets.

Récupération: Modern PDF repair utilities (e.g. qpdf --repair or File2File repair) rescan the file linearly, rebuild the object dictionary, and regenerate a new xref table.

⚠️ PDF viewer fails to load the document or crashes on the last page.

Cause racine: Truncated download or incomplete disk write missing the '%%EOF' marker and trailer dictionary.

Récupération: Inspect file tail using 'tail -c 128 document.pdf'. If missing %%EOF, verify if the root catalog object exists and append an artificial xref table pointing to the highest valid object number.

⚠️ Embedded fonts render as illegible symbols, blank blocks (.notdef), or question marks.

Cause racine: Non-embedded subset fonts referencing system fonts that are missing on the viewing machine, or corrupted ToUnicode CMap tables.

Récupération: Re-embed missing font subsets via Ghostscript (gs -o fixed.pdf -sDEVICE=pdfwrite -dEmbedAllFonts=true broken.pdf) or convert text elements to vector curves.

Analyse de sécurité et vecteurs d'attaque de l'analyseur

PDF is one of the most complex formats in computing, featuring active scripting, embedded files, and dynamic streams. Modern browsers sandbox PDF renderers (e.g. Chrome PDFium) in isolated processes.

Vecteurs d'attaque connus

  • Embedded Acrobat JavaScript (JS API) executing arbitrary code or phishing exploits upon document open.
  • Launch Actions (/Launch) configured to execute external system binaries or shell scripts.
  • Font rendering vulnerabilities in legacy Type 1 and TrueType font parsing engines triggering buffer overflows.
  • Server-Side Request Forgery (SSRF) and NTLM hash credential theft via malicious URI actions (/URI) or form submissions.

Bonnes pratiques de défense: Never open untrusted PDFs with Adobe Acrobat JavaScript execution enabled. For web viewing, always render via browser sandboxed WebAssembly/PDF.js engines, or convert untrusted PDFs to sanitized PDF/A-1b documents.

Origines historiques et jalons

2020ISO 32000-2:2020 released, formalizing modern PDF 2.0 enhancements.
2008Adobe transfers PDF specifications to the International Organization for Standardization as ISO 32000-1.
1993PDF 1.0 launched alongside Adobe Acrobat, supporting text, links, and fonts.
1991Dr. John Warnock initiates the Camelot Project at Adobe to solve digital document portability.

Principaux avantages

  • 100% universal visual fidelity across all devices, operating systems, and printers.
  • Supports digital cryptographic signatures, redaction annotations, and DRM access permissions.
  • Combines vector graphics, scalable fonts, and compressed raster streams into a single self-contained container.

Limites techniques

  • Fixed-layout page canvas makes automatic responsive reflowing on small smartphones challenging.
  • Complex internal object stream hierarchy can lead to bloated file sizes without compaction and unreferenced object elimination.
  • Embedded JavaScript and dynamic launch action capabilities present an expansive attack surface.

Anecdotes techniques

  • The original Camelot project paper envisioned printing documents electronically from any application directly to any display.
  • Every PDF file ends with the exact string '%%EOF' marking the end-of-file trailer.
  • Federal courts and national archives globally mandate PDF/A (ISO 19005) for long-term legal preservation.

Foire aux questions techniques

Why do PDF files look identical on every screen and printer?

PDF specifies absolute Cartesian coordinates (points at 1/72 of an inch) for every glyph, vector curve, and raster image on a fixed page canvas. Because fonts, vector paths, and color profiles are embedded directly in the container, rendering does not depend on local system fonts or printer drivers.

What is the difference between standard PDF and PDF/A?

PDF/A (ISO 19005) is a specialized subset designed for long-term archiving. It strictly bans features that could compromise future readability, such as JavaScript, external font linking, encryption, sound/video embedding, and device-dependent color spaces. All fonts and color profiles must be 100% embedded.

How does PDF compression work in modern documents?

PDF uses multiple discrete compression filters tailored to content types. Text, vector commands, and font metrics are compressed using FlateDecode (zlib/deflate). Photographic images use DCTDecode (JPEG) or JPXDecode (JPEG 2000), while bi-level monochrome scanned text uses JBIG2 or CCITT Fax Group 4 compression.

Why does a PDF file often stay large even after deleting pages or images?

When modifying a PDF, many editors perform an 'incremental update'—they append new object streams and a new xref table to the end of the file without deleting the old objects, allowing fast undo. To actually reduce the file size, the document must undergo an optimization pass that garbage-collects unreferenced objects.