About this role
Role Overview
Build and operate scalable data pipelines that turn raw, structured, and unstructured information into reliable, high quality datasets used for research, analytics, and AI and machine-learning model development. This full-time remote role sits on the core data engineering team and contributes directly to preparing data that powers frontier AI work.
Key Responsibilities- Design, develop, and maintain scalable ETL pipelines for both structured and unstructured data.
- Collect, clean, transform, and organize datasets to support research and ML workflows.
- Perform exploratory data analysis to surface trends, patterns, anomalies, and data quality issues.
- Collaborate with researchers, data scientists, and engineers to prepare datasets for AI and machine-learning initiatives.
- Develop and maintain data models, database schemas, and storage solutions.
- Write and optimize SQL queries for extraction, transformation, and analysis.
- Ensure data quality, integrity, consistency, and security across pipelines and databases.
- Automate data validation, reporting, and recurring data processing workflows.
- Document data pipelines, processes, and technical decisions.
- Troubleshoot pipeline failures and resolve data-related issues.
Required
- Strong proficiency in Python and SQL.
- Hands-on experience designing and maintaining ETL pipelines.
- Experience conducting exploratory data analysis.
- Proficiency with Python data-processing libraries such as Pandas and NumPy.
- Experience with relational databases, particularly PostgreSQL and MySQL.
- Strong understanding of data modeling, database schemas, and data transformation.
- Experience working with both structured and unstructured datasets and ensuring data accuracy and reliability.
- Familiarity with development environments such as Jupyter Notebook, VS Code, or PyCharm.
- Strong analytical, problem-solving, and communication skills.
Nice to have
- Exposure to AI and machine-learning workflows, including scikit-learn.
- Experience with Hugging Face Transformers and familiarity with the OpenAI API or similar AI platforms.
- Experience preparing and managing datasets specifically for AI and ML model development.
- Previous collaboration experience with researchers or data scientists.
- Employment type: Full-time.
- Location: Remote, remote-first workforce.
- Team: Core Data Engineer team.
- Listed compensation for this role: $140, 000 - $180, 000 per year.
- Company pay notice: national pay range for this full-time position is base salary $100, 000 - $150, 000 USD.
- All employees are eligible for equity compensation. Employees may also receive performance-based bonuses depending on role and company policy.
- Benefits include up to 100% reimbursement for health-insurance premiums, paid time off, a 401(k) plan with company match, and additional benefits for a remote-first team.
- The company states it is an equal opportunity employer.
- All employees are eligible for equity and may be eligible for performance-based bonuses, subject to role and company policies.