Sample Stata File
A downloadable sample Stata .dta file with 30 fictional employee records in the modern version 118 format — ready for pandas, R haven, and Stata itself.
sample-employees.dta
Data Preview
| 1 | name | department | role | salary | start_date | office | |
| 2 | Marcus Chen | [email protected] | Engineering | Senior Software Engineer | 155000 | 2019-03-15 | San Francisco |
| 3 | Priya Sharma | [email protected] | Engineering | Staff Engineer | 178000 | 2019-06-01 | San Francisco |
| 4 | David Kim | [email protected] | Engineering | Software Engineer | 125000 | 2021-01-10 | New York |
| 5 | Rachel Torres | [email protected] | Engineering | Engineering Manager | 168000 | 2020-02-20 | San Francisco |
| 6 | James Okafor | [email protected] | Engineering | Junior Developer | 92000 | 2024-06-15 | Austin |
| 7 | Lena Vogt | [email protected] | Engineering | DevOps Engineer | 140000 | 2022-04-01 | New York |
| 8 | Amir Patel | [email protected] | Engineering | Backend Engineer | 132000 | 2023-01-09 | London |
| 9 | Sofia Lindberg | [email protected] | Design | Lead Designer | 145000 | 2019-09-12 | New York |
| 10 | Carlos Rivera | [email protected] | Design | UX Designer | 112000 | 2021-07-20 | San Francisco |
| 11 | Hannah Becker | [email protected] | Design | UI Designer | 105000 | 2022-11-01 | London |
| 12 | Yuki Tanaka | [email protected] | Design | Product Designer | 118000 | 2023-03-14 | San Francisco |
| 13 | Olivia Martin | [email protected] | Marketing | VP of Marketing | 165000 | 2019-04-22 | New York |
| 14 | Ethan Brooks | [email protected] | Marketing | Content Strategist | 95000 | 2021-10-05 | Austin |
| 15 | Nina Kowalski | [email protected] | Marketing | SEO Specialist | 88000 | 2022-08-15 | New York |
| 16 | Daniel Ochoa | [email protected] | Marketing | Marketing Analyst | 91000 | 2023-05-20 | Austin |
| 17 | Samira Hassan | [email protected] | Marketing | Social Media Manager | 82000 | 2024-01-08 | London |
| 18 | Tyler Washington | [email protected] | Sales | Sales Director | 158000 | 2019-11-30 | New York |
| 19 | Jessica Huang | [email protected] | Sales | Account Executive | 110000 | 2020-06-14 | San Francisco |
| 20 | Ryan O'Brien | [email protected] | Sales | Account Executive | 105000 | 2021-03-22 | London |
| 21 | Fatima Al-Rashid | [email protected] | Sales | Sales Development Rep | 72000 | 2023-09-01 | Austin |
| 22 | Kevin Dupont | [email protected] | Sales | Solutions Engineer | 135000 | 2022-01-17 | San Francisco |
| 23 | Megan Stewart | [email protected] | Sales | Account Manager | 98000 | 2024-03-11 | New York |
| 24 | Laura Chen | [email protected] | HR | HR Director | 148000 | 2019-08-05 | New York |
| 25 | Brian Nakamura | [email protected] | HR | HR Business Partner | 105000 | 2020-12-01 | San Francisco |
| 26 | Chloe Dubois | [email protected] | HR | Recruiter | 78000 | 2022-05-23 | London |
| 27 | Angela Moretti | [email protected] | HR | People Operations | 85000 | 2023-07-10 | Austin |
| 28 | Isaac Fernandez | [email protected] | Engineering | Frontend Engineer | 128000 | 2022-09-19 | New York |
| 29 | Sarah Mitchell | [email protected] | Design | Design Systems Lead | 138000 | 2020-04-06 | San Francisco |
| 30 | Omar Farah | [email protected] | Engineering | QA Engineer | 95000 | 2024-02-12 | London |
| 31 | Natalie Park | [email protected] | Marketing | Growth Manager | 108000 | 2021-11-28 | San Francisco |
Schema
| Field | Type | Description |
|---|---|---|
| name | string | Employee full name. |
| string | Work email address on the fictional example.com domain. | |
| department | string | One of Engineering, Sales, Marketing, HR, or Design. |
| role | string | Job title within the department. |
| salary | int32 | Annual salary in USD, stored as a Stata long (32-bit integer). |
| start_date | string | Hire date as an ISO 8601 string (YYYY-MM-DD), not a Stata date value. |
| office | string | Office location — San Francisco, New York, Austin, or London. |
About the Stata Format
Stata’s .dta is the native dataset format of the Stata statistical package, ubiquitous in economics, epidemiology, and survey research. This sample uses format version 118 — the dialect introduced with Stata 14 — which is the one to standardize on: it stores strings as UTF-8 (older versions were effectively Latin-1) and supports large observation counts. It was written by pandas.DataFrame.to_stata(..., version=118).
The format is binary and self-describing: a header carries the observation count and variable metadata, followed by fixed-layout data records. Two Stata-specific conventions show up even in a file this small:
- Variable name rules. Stata names are limited to 32 characters, may contain only letters, digits, and underscores, and cannot start with a digit. The employee headers (
name,start_date, …) already comply, so no renaming occurred — but a CSV header like2024 salaryorfirst-namewould be rewritten on export, which is a common source of column-name drift between Stata and everything else. - Typed columns.
salaryis stored as a Statalong(32-bit integer), and each string column is a fixed-widthstr#type sized to its longest value. Unlike CSV, the types travel with the file.start_dateis kept as an ISO 8601 string rather than a Stata internal date (days since 1960-01-01) so it reads identically in pandas, R, and Stata without format directives.
What this file omits is also informative: real research .dta files usually carry variable labels, value labels (integer codes mapped to text), and missing-value sentinels (., .a–.z). This sample has none, making it a clean baseline before testing those features with your own data.
To open it, use pandas.read_stata(), R’s haven::read_dta(), or Stata 14+ — or follow the DTA to CSV guide for every free conversion route, including no-code options. The Parquet and Feather siblings have in-browser converters if you need CSV output today.