The Challenge of High-Volume Recruitment
When a role attracts a large number of applications, reviewing each CV manually quickly becomes a bottleneck. The problem is not just the time spent reading - it is the data entry that follows. For every CV that arrives by email, from an agency, or collected at a job fair, someone has to open the document, read the contact details, type them into the ATS, attach the file, and move on to the next one.
At any meaningful scale, this creates real operational problems:
- Manual entry errors - a transposed digit in a phone number or a misspelled email means a candidate you cannot reach
- Delays in outreach - candidates who applied on Monday may not be in your system until Thursday, by which time they may have accepted another offer
- Administrative burden on recruiters - time spent on data entry is time not spent evaluating candidates or building relationships
- Risk of lost applications - CVs that sit in an inbox or on a desktop for days can slip through the cracks entirely
Bulk CV upload addresses this directly. Instead of processing applications one at a time, the recruiter selects a batch of files, uploads them together, and the ATS creates candidate profiles automatically - with data extracted from each document.
What Bulk CV Upload Actually Does
Uploading Multiple Files at Once
Most modern ATS platforms allow you to select a batch of CV files - PDFs, Word documents, sometimes RTF - and upload them in a single action. The system processes each file and creates a separate candidate profile for each one. The recruiter does not need to open, read, or manually enter data from any individual document.
Treegarden supports uploading up to 50 CVs per batch, with a maximum of 20 MB per file. Files are processed sequentially with real-time feedback. ZIP archives are not accepted for security reasons - each document must be uploaded directly as a file.
Automated CV Parsing
Parsing is the process by which the ATS reads an uploaded document and extracts structured data from it. The goal is to turn unstructured text - a CV written by a human, in their own format - into a set of searchable, filterable fields: candidate name, email address, phone number, work history, education, and skills.
Modern parsers handle standard formats well. A straightforward chronological CV in PDF or Word will typically yield accurate extractions for name, contact details, education institutions, job titles, and employer names. Skills extraction is less consistent - candidates describe competencies in many different ways, and parsers vary in how well they interpret this.
Parsing accuracy drops with certain CV designs. Multi-column layouts, heavily graphical resumes, tables used for structure, and scanned image PDFs all present challenges. The parser reads text in a linear sequence; a two-column layout can cause it to interleave content from both columns, producing garbled output. Scanned PDFs - where the content is an image rather than selectable text - require optical character recognition, which introduces additional error. The practical guidance: wherever possible, ask candidates to submit CVs in PDF or Word format with a straightforward layout.
Even when parsing partially fails, a well-designed ATS will still create a candidate profile with the original document attached, so the recruiter can complete the missing fields manually. No application should be lost because the parser encountered an unusual format.
From Upload to Profile
The typical flow after a bulk upload works as follows:
- The recruiter selects a batch of CV files and initiates the upload
- The ATS processes each document, extracting name, contact details, work history, education, and skills
- A candidate profile is created for each file, with the original document attached
- The recruiter reviews the extracted data, corrects any parsing errors, and enriches profiles as needed
- Profiles can be assigned to a specific vacancy, added to the talent pool, or tagged by source and date
The review step is important. A light spot-check of 5-10 records after any bulk upload - verifying that names, emails, and phone numbers have parsed correctly - catches the most common errors before they cause problems downstream.
What CV Formats Are Supported?
PDF and Word (.docx) are the most reliably parsed formats across ATS platforms. Plain text files also parse cleanly. Some platforms additionally support older Word formats (.doc) and RTF.
Formats that parse less accurately include:
| Format | Parsing reliability | Notes |
|---|---|---|
| PDF (text-based, simple layout) | High | Best general-purpose format |
| Word .docx | High | Standard for many candidates |
| Plain text (.txt) | High | Clean but uncommon |
| PDF (multi-column or graphic-heavy) | Medium | Column layout disrupts linear reading |
| Scanned image PDF | Low | Requires OCR; significant error rate |
| HTML, Google Docs link, image files | Not supported | Convert to PDF or Word before uploading |
When setting submission guidelines for candidates - or when specifying requirements to recruitment agencies - asking for PDF or Word in a single-column layout will give you the best parsing results.
Duplicate Candidate Detection
Candidates apply through multiple channels. The same person may send a CV directly by email, apply through a job board, and be submitted by a recruitment agency - all for the same role, or across different roles over time. Without duplicate detection, each submission creates a separate profile, resulting in a cluttered database, inaccurate reporting, and the risk of contacting the same candidate twice through different touchpoints.
A well-designed ATS detects duplicates during the upload process, typically by matching on email address, phone number, or a combination of name and contact details. When a match is found, the system presents the recruiter with a choice: merge the new record with the existing profile (combining application history and notes), or keep them as separate records (useful when tracking two genuinely distinct applications independently).
Merging is the right call in most cases - it keeps the candidate database clean and ensures that all interactions with a candidate are visible in one place. A clean database is also important for accurate recruitment reporting: if one candidate has three profiles, your application counts and conversion rates will be skewed.
Configuring duplicate detection before running your first bulk upload, rather than cleaning up afterwards, is significantly less work.
GDPR Considerations for Bulk CV Processing
When CVs arrive via bulk upload - whether from a job fair, an agency batch, or a LinkedIn export - the same data protection obligations apply as for any other candidate data. A few considerations specific to bulk processing are worth flagging.
Lawful basis. You need a lawful basis for processing each candidate's personal data. For active recruitment, "legitimate interests" can apply if the candidate submitted their CV in response to a vacancy or in anticipation of future opportunities. Consent may be more appropriate in some contexts. The key point is that you should not bulk-upload CVs without a clear recruitment purpose and an applicable lawful basis for every individual in the batch.
Privacy notice. Candidates whose CVs you process should receive, or have received, a privacy notice explaining how their data is used and stored. If you receive a batch of CVs from an agency, check what privacy notice the agency has already provided to candidates.
Retention periods. Set a retention reminder when you upload a batch. CVs for unsuccessful candidates should not be kept indefinitely. Many organisations use a 6-12 month retention period for active talent pool members and shorter periods for candidates who do not meet minimum criteria.
Right to deletion. Any candidate can request deletion of their data. Your ATS should make it straightforward to locate and delete a candidate's profile and all associated documents.
Best Practices for Bulk CV Processing
The following checklist covers the most important steps for a clean, compliant bulk upload process:
- Spot-check parsed data after upload - review 5-10 records to verify that names, email addresses, and phone numbers have been extracted correctly before taking further action
- Tag the batch with source, date, and role - knowing that a group of profiles came from a specific job fair or agency on a specific date makes them easier to manage and report on
- Set a retention reminder at upload time - decide upfront how long these profiles should be kept and set a reminder to review or delete them at that point
- Review and enrich profiles before sharing with hiring managers - parsed profiles may have gaps; a brief enrichment pass ensures hiring managers receive accurate, complete information
- Configure duplicate detection before your first bulk upload - clean deduplication at import is far less effort than database cleanup after the fact
- Ask candidates and agencies to submit in PDF or Word - standardising the submission format significantly improves parsing accuracy
- Archive or delete profiles that clearly do not meet minimum criteria promptly - keeping only relevant profiles makes the database more useful and supports GDPR compliance
Treegarden and Bulk CV Upload
Treegarden supports uploading up to 50 CVs per batch with automated parsing for name, contact details, work history, education, and skills. Parsed profiles are immediately searchable and can be assigned to a specific vacancy or added to your talent pool. Duplicate detection is built in and runs automatically during each upload.
Where parsing cannot extract all data - for example, from an unusual CV layout - the original document is still stored and the profile is created with the available information, so nothing is lost. AI scoring in Treegarden is advisory: it helps you prioritise which profiles to review first but does not auto-reject any candidate. Every hiring decision remains with the recruiter.
Book a demo to see bulk CV processing in action.
Frequently Asked Questions
How accurate is automated CV parsing?
Accuracy varies by format and CV design. Standard chronological CVs in PDF or Word format typically parse well - name, email, phone, education, and job titles are extracted reliably. Skills extraction is less consistent because candidates describe them in many ways. Heavily formatted CVs (with columns, tables, or graphics) and scanned image PDFs are the most error-prone because the parser reads text left-to-right and the layout disrupts this. Most ATS systems show confidence indicators or flag uncertain fields for manual review. Plan for a light review pass after any bulk upload, especially for senior or specialist roles where data accuracy matters most.
Can an ATS handle CVs in different formats?
Most modern ATS platforms support PDF, Word (.docx), and plain text. Some also handle older formats like .doc or RTF. The most reliably parsed format is a well-structured PDF or Word document. If you receive CVs in formats the ATS does not support (e.g. Google Docs link, HTML, or image files), you may need to convert them first. For job fair or agency submissions where you receive a mix of formats, it is worth setting a clear submission guideline asking for PDF or Word to maximise parsing accuracy.
What happens when a candidate appears twice in the system?
Modern ATS platforms detect duplicates based on matching email address, phone number, or name combinations. When a duplicate is detected, the system typically presents you with a choice: merge the two records (combining their application history and notes) or keep them separate (useful if the candidate is applying for genuinely different roles at different times). Merging is usually the right call for the same person; keeping separate is only useful when you want to track two distinct applications independently. A clean database without duplicates is essential for accurate reporting and avoiding awkward double-outreach to candidates.
What is the difference between CV parsing and CV screening?
Parsing is a data extraction process: the system reads a CV and converts unstructured text into structured fields (name, contact, work history, skills). Screening is an evaluation process: a human or AI tool assesses whether the parsed data meets the criteria for a specific role. In practice, parsing happens first and enables screening to be done at scale - once CVs are structured, you can filter and sort by specific criteria. Automated screening (AI matching scores) should always be treated as a starting point for human review, not a final decision. Treegarden's AI scoring is advisory: it helps prioritise review order but never auto-rejects candidates.