Skip to content

Photo Metadata and Privacy

Document Metadata: What PDFs, Word and Excel Files Record

Every document you send carries a hidden record of who made it, what software they used and sometimes what they deleted. Here is what lives where, and how to clear it before the file leaves your hands.

·Creator of ToolFiddle··16 min read

You send a quote as a PDF. The client opens the properties panel and sees your full name, the fact that it came out of a word processor you would rather they did not know you use, and a modification timestamp of 23:41 on a Sunday. Nothing damaging. Slightly more than you meant to say.

Then there is the version that does damage. A spreadsheet where the pivot table still holds every row of a source sheet you deleted. A contract where tracked changes reveal the clause your side wanted and gave up on. A report where the redaction is a black rectangle and the text underneath copies out cleanly.

Document metadata in PDF, Word and Excel files works much like the metadata in a photo, with one important difference: office documents can retain content you thought you removed, not just facts about the file. That is the part worth taking seriously.

This post covers where each format stores its metadata, which leaks are genuinely dangerous rather than merely untidy, and how to clean each one properly before the file leaves your hands.

The short answer

PDFs store metadata in two places: a document information dictionary with the keys Title, Author, Subject, Keywords, Creator, Producer, CreationDate and ModDate, and usually an XMP packet carrying a second copy in XML. Office files are ZIP archives of XML, with author, last-modified-by, revision count, total editing time, company and template path sitting in docProps/core.xml and docProps/app.xml.

The dangerous leaks are not those fields. They are retained content: tracked changes, comments, hidden rows and columns, hidden worksheets, pivot caches holding data whose source you deleted, images cropped in the document but stored whole, and speaker notes in presentations. Add the redaction trap, where a black box leaves selectable text beneath it.

Clean Office files with the Document Inspector before exporting. Clean PDFs by removing hidden information rather than trusting a visual check.

What a PDF stores about you

Every PDF can carry a document information dictionary. It is a small set of key-value pairs, and any reader can display them.

  • Title, Subject and Keywords, usually empty unless someone filled them in
  • Author, which very often holds a real name pulled from the account that created the source file
  • Creator, meaning the application the content originated in, such as Microsoft Word or LaTeX
  • Producer, meaning the library or driver that actually wrote the PDF bytes
  • CreationDate and ModDate, in the PDF date format D:YYYYMMDDHHmmSSOHH'mm', so a file made at 09:45 on 15 March 2026 in British Summer Time reads D:20260315094512+01'00'

The Creator and Producer pair is more revealing than people expect. Producer strings are distinctive: a file made through the macOS print pipeline names the Quartz PDF context and the OS version, one printed from Chrome names Skia, one from LibreOffice names LibreOffice and its version number. None of that is secret. It does describe your working setup precisely, and it links every document you produce on that machine to every other one.

Alongside the dictionary sits an XMP packet, the same XML container photographs use, holding dc:title, dc:creator, pdf:Producer, xmp:CreateDate, xmp:ModifyDate and xmpMM:DocumentID. Clearing one copy and not the other is the most common mistake in PDF cleaning. The breakdown of how XMP works covers the container itself, since it behaves identically here.

Beyond metadata, a PDF can carry embedded file attachments, JavaScript, form field values that a recipient can read even when fields look empty, annotation and comment threads, bookmarks, and named destinations that sometimes preserve old section headings. It can also carry its own history. PDFs support incremental updates, where a save appends new objects and a new cross-reference table rather than rewriting the file. Earlier versions of a page can therefore still be present in the bytes. A file that has grown from 240 KB to 1.1 MB across three rounds of small edits is worth a second look.

Office files are ZIP archives, and you can look inside

This is the single most useful thing to know about modern Office formats. A .docx, .xlsx or .pptx is an Open Packaging Conventions package, which is a ZIP file with a defined folder structure. Copy one, change the extension to .zip, and open it.

Inside, the metadata lives in a docProps folder.

docProps/core.xml holds the properties most people mean by document properties:

  • dc:creator, the original author
  • cp:lastModifiedBy, whoever saved it most recently
  • cp:revision, a save counter
  • dcterms:created and dcterms:modified, timestamps
  • dc:title, dc:subject, cp:keywords, dc:description, cp:category, cp:contentStatus
  • cp:lastPrinted, which quietly records that a document was printed and when

docProps/app.xml holds application-level detail that surprises people:

  • Application and AppVersion
  • Company and Manager, often inherited from a template or an organisational install
  • Template, the path or name of the template the document was created from
  • TotalTime, cumulative editing minutes
  • Pages, Words, Characters, Paragraphs, Lines
  • TitlesOfParts, which for a spreadsheet lists every worksheet name, including hidden ones, and for a presentation lists every slide title

There may also be docProps/custom.xml, which is where document management systems stamp matter numbers, client codes, classification labels and review identifiers. In organisations that use such systems, this file is frequently the most sensitive thing in the package.

The main content parts matter too. In a Word file, word/document.xml is the body, word/comments.xml holds comment text with author names and initials, word/people.xml can hold contributor identity information, and word/settings.xml contains revision save identifiers. Those identifiers are used in document forensics to determine whether two files came from the same editing session, which is an interesting property to be aware of even if it rarely affects you.

The leaks that actually matter

Fields are tidy problems. Retained content is the one that causes trouble, because the information is not a fact about the file, it is material somebody believed they had deleted.

Tracked changes left switched on

Accept-all-changes is not the same as turning tracking off, and neither is setting the view to Simple Markup. A document that visually shows only final text can still hold every insertion and deletion, since deleted text lives in w:del elements inside the XML and is recoverable in full. Negotiated documents are the obvious risk: the price you started at, the indemnity clause you dropped, the paragraph you softened after legal review.

Comments nobody removed

Resolving a comment does not delete it. Hiding the review pane does not delete it. Comments travel with the file, carry author names, and often contain the most candid writing in the whole document, because people write comments assuming an internal audience.

Hidden rows, columns and worksheets

Hiding is a display setting. The data stays in the sheet XML with a hidden attribute, and unhiding takes one right-click. Worksheets have a third state, veryHidden, which does not appear in the normal unhide dialog and can only be changed through the VBA editor. That one has a habit of being used deliberately and then forgotten by whoever inherits the file.

Pivot caches holding deleted source data

This is the leak I would put first if I could only warn about one. A pivot table stores its own copy of the source records inside the workbook so it can recalculate without touching the source. Delete the source worksheet and the cache normally survives in xl/pivotCache/. Anyone can double-click a value in the pivot table and the software will helpfully generate a new sheet containing the underlying rows. Salary tables, customer lists and unredacted transaction records have all been shared this way by people who did exactly what they thought was right.

Cropped images that are not really cropped

Cropping a picture in Word, Excel or PowerPoint changes how much of it is displayed. The full original image usually remains in the media folder inside the package. Crop a screenshot to hide a column, or crop a photograph to remove a person, and the complete version is still there for anyone who unzips the file.

Speaker notes and off-canvas objects

Notes live in ppt/notesSlides/ and export to PDF when the wrong option is chosen. Objects dragged just outside the slide boundary are still in the file and still in the XML, even though nobody sees them in presentation mode. Slides hidden from a presentation remain in the deck.

External data connections, linked workbooks and hyperlinks can expose internal server paths, network share names and file structures. Those paths tell a reader how your organisation is arranged, which is a small thing that occasionally matters a great deal.

The redaction trap

A black rectangle drawn over text is not redaction. It is a drawing object placed above the page content, and the characters underneath remain in the content stream exactly as they were. Select the area, copy, paste into a text editor, and the text appears. Any extraction tool will do the same automatically.

The same failure comes in several disguises. Highlighting text in black leaves the text. Setting the font colour to white leaves the text. Placing an image over a paragraph leaves the paragraph. Exporting to PDF from a document where you covered something with a shape carries both the shape and the covered content into the PDF.

Real redaction has two steps, and both are required. First the underlying characters and image regions have to be deleted from the content, not merely obscured. Then the file has to be written out so that no earlier version remains. Acrobat Pro’s redaction tool does this when you apply the marks, and its Remove Hidden Information command handles the metadata that redaction alone does not touch, since a document titled “Complaint against J Whitfield 2026” is not much helped by blacking out the name inside it.

If you lack a proper redaction tool, the reliable fallback is rasterising. Cover the text, export each page as an image, then rebuild the document from those images. It is crude, it destroys searchable text and accessibility, and it produces a bigger file. It also works, because an image of a page contains no character data. Tools like PDF to Image and Image to PDF will do both halves of that round trip in the browser without the pages being uploaded anywhere.

One caveat, stated plainly. If you are redacting for a court, a regulator, a freedom of information response or any formal disclosure, follow the procedure your jurisdiction or organisation specifies rather than a general method from an article. Requirements differ, they change, and the consequences of getting it wrong in those settings are not the same as sending a slightly chatty PDF to a client.

Where the metadata lives, format by format

Format Metadata location Retained content risks How to clear it
PDF Document info dictionary plus an XMP packet Attachments, JavaScript, form values, comments, incremental update history Remove hidden information in a full PDF editor, or rasterise and rebuild
DOCX docProps/core.xml, app.xml, custom.xml Tracked changes, comments, hidden text, cropped images, embedded objects Document Inspector, then re-check properties
XLSX Same docProps files Hidden rows, columns and sheets, pivot caches, defined names, external links Document Inspector, delete pivot caches, unhide and check everything first
PPTX Same docProps files Speaker notes, hidden slides, off-slide objects, cropped images Document Inspector, plus compress pictures to delete cropped areas
DOC, XLS, PPT (legacy) Binary summary information streams Older save behaviour could retain deleted text in the file Convert to the modern format, inspect, then export fresh
ODT, ODS, ODP meta.xml inside the ZIP package Tracked changes, comments, hidden content Remove personal information in the document settings, then check meta.xml
RTF Embedded property groups Tracked changes can survive conversion Convert to DOCX, inspect, export again
Plain text and CSV None inside the file Filesystem timestamps only Nothing to strip inside the file itself

The last row is worth a moment. If all you need to send is data, CSV carries no metadata, no formulas, no hidden sheets and no cache. Choosing a simpler format is often faster than cleaning a complex one.

A worked example: the CV that names someone else

Say Rafael is applying for graduate roles. He downloads a CV template, edits it thoroughly, and sends it to 14 employers over three weeks.

Unzip one of those files and docProps/app.xml reports Template: Modern_Chronological_v4.dotx, Company: Redwood Careers Ltd, and TotalTime: 63. docProps/core.xml reports dc:creator: J. Whitfield, cp:lastModifiedBy: Rafael, and cp:revision: 22.

Read that as a recruiter would. The document is authored by a person Rafael has never met, associated with a careers company he has never worked for, and it states that the total editing time across all sessions was 63 minutes. Rafael’s covering email describes the CV as tailored to the role. Twenty-two saves and 63 minutes is not a scandal, but it is a detail he did not choose to disclose and cannot explain if asked.

The arithmetic on fixing it favours doing it early. Running the Document Inspector on the master file once, before he starts making per-employer copies, takes roughly 90 seconds. Cleaning 14 already-sent variants afterwards takes about 40 seconds each, so 14 times 40 seconds is 560 seconds, or 9 minutes 20 seconds, and by then the originals have already been read. One pass on the master, before the first copy exists, is the whole fix.

If any of those applications included a portfolio image or a scanned certificate, the same logic applies to those files. Photos carry their own metadata block, and the EXIF Viewer will show you what a scan of a document is carrying in about ten seconds. Everything runs locally in your browser, which matters when the file you are checking is a signed contract or a passport scan you would rather not upload anywhere.

Cleaning each format, step by step

Work on a copy. Every routine below removes things permanently, and the undo history does not survive a save.

For a Word, Excel or PowerPoint file on Windows:

  1. Save a copy of the file under a new name. This copy is the one you will clean and send.
  2. Turn off track changes, then accept or reject every remaining change explicitly rather than relying on the view setting.
  3. Delete all comments through the review tools, including resolved ones.
  4. In a spreadsheet, unhide every row, column and worksheet so you can see what is actually there. Check for very hidden sheets through the VBA editor if the workbook came from elsewhere.
  5. Select any pivot table, open its options, and clear the retained cache or switch off saving source data with the file.
  6. Select cropped images, then use Compress Pictures with the option to delete cropped areas.
  7. Go to File, then Info, then Check for Issues, then Inspect Document. Run it and review every category it flags before removing.
  8. Reopen File then Info and confirm the properties panel is now empty of names.
  9. Export to PDF, then open the PDF properties and confirm the Author and Title fields did not come across.

On macOS, the Document Inspector is not present in the same form. The nearest equivalent is the preference to remove personal information from the file on save, which covers author names but not tracked changes, comments or hidden content. Do those steps manually and verify by unzipping the file.

For a PDF you did not create yourself, or one you cannot verify:

  1. Open the document properties and clear Title, Author, Subject and Keywords.
  2. Use the editor’s remove hidden information or sanitise command to clear the XMP packet, attachments, scripts, hidden layers and comments together.
  3. Save as a new file rather than saving over the original, so no incremental update chain is preserved.
  4. Reopen the saved file and check the properties again. If the Producer field now names your PDF editor, the file was genuinely rewritten.
  5. If you have no editor capable of that, rasterise: export the pages as images, then rebuild a PDF from them. Nothing textual survives, which is the point.

Running a PDF through PDF Compress rebuilds the file structure as part of re-encoding, which removes stale incremental-update layers as a side effect and usually shrinks the file. Treat that as a helpful consequence rather than a redaction method. For clearing metadata across mixed batches of images and documents, Metadata Remover handles the general case in the browser, with nothing uploaded.

Things that go wrong at this stage

Exporting to PDF as a cleaning step is the most common misunderstanding. Word copies its title and author straight into the PDF information dictionary, and depending on the export options it can carry comments across as PDF annotations and speaker notes onto the page. Conversion changes the container, not the contents.

Renaming does even less. A file called final_v9_CLEAN.docx contains exactly what draft_v1.docx contained, because the name is a filesystem label and the metadata is inside the bytes.

Running the Document Inspector too aggressively causes the opposite problem. It will strip custom document properties that your document management system depends on, remove headers and footers containing required legal wording, and delete ink annotations somebody wanted. Read the list before you click remove, on a copy, every time.

Legacy formats carry an old and genuinely serious failure mode. Historic word processor save behaviour could append changes rather than rewriting the document, leaving previously deleted text sitting in the file. If you receive an old .doc or .xls, do not simply forward it. Convert it, inspect it, and export a fresh copy.

Collaborative cloud documents shift the problem rather than removing it. Version history, comment threads and edit attribution live on the server. Downloading a copy usually leaves that history behind, but sharing a link does not, and a link shared with the wrong permission level gives a recipient everything the document has ever been. Decide deliberately between sending a file and sending a link.

And there is the small one that catches almost everybody: checking the file you cleaned rather than the file you attached. Clean the copy, then open the attachment from the sent folder and look at its properties. The gap between those two files is where mistakes live.

A short pre-send routine

For anything going outside your organisation, the whole check takes under two minutes.

Open the properties panel and read what a recipient would read. If your name, your employer, a client name or a template path is in there and does not need to be, clear it. Unhide everything in a spreadsheet and look. Confirm that anything you meant to remove has actually gone rather than being covered. Then send, and check the attachment rather than the source.

That routine is worth running against your own files first, before you need it. Take three documents you have already sent this month, look at what they carried, and you will know within five minutes whether this is a real problem for you or an occasional tidiness issue. Same principle as checking your own photos before posting, described in the rundown of what platforms strip from images, and the file inspector is a reasonable place to start with any image attachments in the same batch.

Frequently asked questions

How do I remove metadata from a PDF?

Open the document properties and clear Title, Author, Subject and Keywords, then remove the XMP packet as well, since it holds a second copy of the same values. In Acrobat Pro the Remove Hidden Information command handles both plus attachments, scripts and comments. If you cannot verify the result, rasterise the pages and rebuild the PDF from images, which drops everything including the text layer.

What does the Producer field in a PDF tell someone?

Producer names the library or driver that generated the PDF, and Creator names the application the content came from. Together they usually reveal your operating system, the software you use and often its version. That is rarely sensitive on its own, but it describes your setup precisely and links documents produced on the same machine to each other.

How do I remove my name from a Word document?

Your name sits in docProps/core.xml as dc:creator and cp:lastModifiedBy, and possibly on every tracked change and every comment. On Windows, File then Info then Check for Issues then Inspect Document clears document properties and personal information. On macOS, use the preference to remove personal information on save. Verify afterwards by reopening the file properties rather than assuming.

Is drawing a black box over text in a PDF safe redaction?

No. A rectangle is a drawing object placed above the page content, and the text underneath stays in the content stream where it can be selected, copied or extracted by any tool. Real redaction deletes the underlying characters before flattening the file. If your software has no redaction feature, rasterise the page after covering the text and rebuild the document from the images.

Does a pivot table keep deleted source data?

Often yes. A pivot table caches its own copy of the source records inside the workbook so it can recalculate without the source present. Delete the source sheet and the cache usually survives, and anyone can rebuild the underlying rows by double-clicking a value in the table. Clear the cache, or switch off saving source data with the file, before sharing the workbook.

Does converting a Word file to PDF remove its metadata?

No. Exporting to PDF typically copies the title and author across into the PDF document information dictionary, and it can carry comments and hidden slide notes depending on the settings you use. Conversion changes the container, not the contents. Inspect the document before you export, then check the resulting PDF properties as a separate step.

What to do with this

Most document metadata is harmless. Your name on a quote, a Producer string naming your PDF library, an editing time of 63 minutes: none of that will hurt you, and treating every field as a hazard is how people end up ignoring the whole subject.

The retained content is different. Tracked changes, comments, hidden rows, pivot caches, uncropped images and text under a black box are all cases where a document contains something its author believed had been removed. Those are worth two minutes before anything leaves your hands, and they are worth checking on your own files today rather than after a mistake.

Pick one document you sent recently. Unzip it, read docProps, look at what it says. If nothing surprises you, the risk here is low for the way you work. If something does, you now know exactly which routine to run.

Frequently asked questions

How do I remove metadata from a PDF?

Open the document properties and clear Title, Author, Subject and Keywords, then remove the XMP packet as well, since it holds a second copy of the same values. In Acrobat Pro the Remove Hidden Information command handles both plus attachments, scripts and comments. If you cannot verify the result, rasterise the pages and rebuild the PDF from images, which drops everything including the text layer.

What does the Producer field in a PDF tell someone?

Producer names the library or driver that generated the PDF, and Creator names the application the content came from. Together they usually reveal your operating system, the software you use and often its version. That is rarely sensitive on its own, but it describes your setup precisely and links documents produced on the same machine to each other.

How do I remove my name from a Word document?

Your name sits in docProps/core.xml as dc:creator and cp:lastModifiedBy, and possibly on every tracked change and every comment. On Windows, File then Info then Check for Issues then Inspect Document clears document properties and personal information. On macOS, use the preference to remove personal information on save. Verify afterwards by reopening the file properties rather than assuming.

Is drawing a black box over text in a PDF safe redaction?

No. A rectangle is a drawing object placed above the page content, and the text underneath stays in the content stream where it can be selected, copied or extracted by any tool. Real redaction deletes the underlying characters before flattening the file. If your software has no redaction feature, rasterise the page after covering the text and rebuild the document from the images.

Does a pivot table keep deleted source data?

Often yes. A pivot table caches its own copy of the source records inside the workbook so it can recalculate without the source present. Delete the source sheet and the cache usually survives, and anyone can rebuild the underlying rows by double-clicking a value in the table. Clear the cache, or switch off saving source data with the file, before sharing the workbook.

Does converting a Word file to PDF remove its metadata?

No. Exporting to PDF typically copies the title and author across into the PDF document information dictionary, and it can carry comments and hidden slide notes depending on the settings you use. Conversion changes the container, not the contents. Inspect the document before you export, then check the resulting PDF properties as a separate step.

Tools that go with this

Same promise: everything runs in your browser, nothing gets uploaded.

Keep reading