Sample Parquet File
A downloadable sample Parquet file with 30 fictional employee records — columnar, snappy-compressed, with an embedded schema. Ready for readers, viewers, and converters.
sample-employees.parquet
Data Preview
| 1 | name | department | role | salary | start_date | office | |
| 2 | Marcus Chen | [email protected] | Engineering | Senior Software Engineer | 155000 | 2019-03-15 | San Francisco |
| 3 | Priya Sharma | [email protected] | Engineering | Staff Engineer | 178000 | 2019-06-01 | San Francisco |
| 4 | David Kim | [email protected] | Engineering | Software Engineer | 125000 | 2021-01-10 | New York |
| 5 | Rachel Torres | [email protected] | Engineering | Engineering Manager | 168000 | 2020-02-20 | San Francisco |
| 6 | James Okafor | [email protected] | Engineering | Junior Developer | 92000 | 2024-06-15 | Austin |
| 7 | Lena Vogt | [email protected] | Engineering | DevOps Engineer | 140000 | 2022-04-01 | New York |
| 8 | Amir Patel | [email protected] | Engineering | Backend Engineer | 132000 | 2023-01-09 | London |
| 9 | Sofia Lindberg | [email protected] | Design | Lead Designer | 145000 | 2019-09-12 | New York |
| 10 | Carlos Rivera | [email protected] | Design | UX Designer | 112000 | 2021-07-20 | San Francisco |
| 11 | Hannah Becker | [email protected] | Design | UI Designer | 105000 | 2022-11-01 | London |
| 12 | Yuki Tanaka | [email protected] | Design | Product Designer | 118000 | 2023-03-14 | San Francisco |
| 13 | Olivia Martin | [email protected] | Marketing | VP of Marketing | 165000 | 2019-04-22 | New York |
| 14 | Ethan Brooks | [email protected] | Marketing | Content Strategist | 95000 | 2021-10-05 | Austin |
| 15 | Nina Kowalski | [email protected] | Marketing | SEO Specialist | 88000 | 2022-08-15 | New York |
| 16 | Daniel Ochoa | [email protected] | Marketing | Marketing Analyst | 91000 | 2023-05-20 | Austin |
| 17 | Samira Hassan | [email protected] | Marketing | Social Media Manager | 82000 | 2024-01-08 | London |
| 18 | Tyler Washington | [email protected] | Sales | Sales Director | 158000 | 2019-11-30 | New York |
| 19 | Jessica Huang | [email protected] | Sales | Account Executive | 110000 | 2020-06-14 | San Francisco |
| 20 | Ryan O'Brien | [email protected] | Sales | Account Executive | 105000 | 2021-03-22 | London |
| 21 | Fatima Al-Rashid | [email protected] | Sales | Sales Development Rep | 72000 | 2023-09-01 | Austin |
| 22 | Kevin Dupont | [email protected] | Sales | Solutions Engineer | 135000 | 2022-01-17 | San Francisco |
| 23 | Megan Stewart | [email protected] | Sales | Account Manager | 98000 | 2024-03-11 | New York |
| 24 | Laura Chen | [email protected] | HR | HR Director | 148000 | 2019-08-05 | New York |
| 25 | Brian Nakamura | [email protected] | HR | HR Business Partner | 105000 | 2020-12-01 | San Francisco |
| 26 | Chloe Dubois | [email protected] | HR | Recruiter | 78000 | 2022-05-23 | London |
| 27 | Angela Moretti | [email protected] | HR | People Operations | 85000 | 2023-07-10 | Austin |
| 28 | Isaac Fernandez | [email protected] | Engineering | Frontend Engineer | 128000 | 2022-09-19 | New York |
| 29 | Sarah Mitchell | [email protected] | Design | Design Systems Lead | 138000 | 2020-04-06 | San Francisco |
| 30 | Omar Farah | [email protected] | Engineering | QA Engineer | 95000 | 2024-02-12 | London |
| 31 | Natalie Park | [email protected] | Marketing | Growth Manager | 108000 | 2021-11-28 | San Francisco |
Schema
| Field | Type | Description |
|---|---|---|
| name | string | Employee full name. |
| string | Work email address on the fictional example.com domain. | |
| department | string | One of Engineering, Sales, Marketing, HR, or Design. |
| role | string | Job title within the department. |
| salary | int32 | Annual salary in USD, stored as a 32-bit integer column. |
| start_date | string | Hire date as an ISO 8601 string (YYYY-MM-DD). |
| office | string | Office location — San Francisco, New York, Austin, or London. |
About the Parquet Format
Parquet is a columnar storage format: instead of writing records one after another the way CSV does, it writes each column’s values together. This sample stores all 30 names, then all 30 emails, and so on — grouped into a single row group, with each column chunk compressed independently using snappy.
That layout is why Parquet dominates analytics workloads:
- Column pruning. A query that needs only
salaryreads only the salary chunk, skipping the other six columns entirely. - Predicate pushdown. Per-chunk statistics (min/max) in the footer let engines skip whole row groups that can’t match a filter.
- Embedded schema. The footer records every column’s name and physical type, so
salarycomes back as anint32— no header-sniffing or type inference, unlike CSV.
The trade-off is that none of this is inspectable with cat. The file is binary, footer-indexed, and only meaningful to a Parquet reader — hence no raw-contents block on this page. Open it with the Parquet Viewer, convert it with Parquet to CSV, or load it in code: pd.read_parquet(...), duckdb.sql("SELECT * FROM 'sample-employees.parquet'"), or arrow::read_parquet() in R.
One honest caveat: at 30 rows, format overhead (footer metadata, dictionary pages) dominates and the file is larger per-row than its CSV twin. Parquet’s advantages compound at thousands to billions of rows; this sample is sized for testing correctness, not benchmarking.
All eleven formats in this section carry the same 30 employee records. The nearest relative is the Feather sample — also columnar and binary, but optimized for interchange speed rather than storage efficiency. To produce files like this from your own data, use CSV to Parquet.