Cohere · 2 days ago
Member of Technical Staff, Pre-Training Data
Cohere is a company focused on scaling intelligence to serve humanity by training and deploying frontier models for AI systems. The role of Member of Technical Staff, Pre-Training Data involves developing data pipelines for advanced language models and conducting data ablations to enhance model performance, directly impacting natural language processing innovations.
Artificial Intelligence (AI)Generative AIMachine LearningNatural Language Processing
Responsibilities
Conduct data ablations to assess data quality and experiment with data mixtures to enhance model performance
Develop robust data modeling techniques to ensure datasets are structured and formatted for optimal training efficiency
Research and implement innovative data curation methods, leveraging Cohere’s infrastructure to drive advancements in natural language processing
Collaborate with cross-functional teams, including researchers and engineers, to ensure data pipelines meet the demands of cutting-edge language models
Qualification
Required
Strong software engineering skills, with proficiency in Python and experience building data pipelines
Familiarity with curriculum learning, data mixing and data attribution
Familiarity with data processing frameworks such as Apache Spark, Apache Beam, Pandas, or similar tools
Experience working with large-scale datasets, including web data, code data, and multilingual corpora
Knowledge of data quality assessment techniques and experimentation with data mixtures
A passion for bridging research and engineering to solve complex data-related challenges in AI model training
Preferred
paper at top-tier venues (such as NeurIPS, ICML, ICLR, AIStats, MLSys, JMLR, AAAI, Nature, COLING, ACL, EMNLP)
Benefits
An open and inclusive culture and work environment
Work closely with a team on the cutting edge of AI research
Weekly lunch stipend, in-office lunches & snacks
Full health and dental benefits, including a separate budget to take care of your mental health
100% Parental Leave top-up for up to 6 months
Personal enrichment benefits towards arts and culture, fitness and well-being, quality time, and workspace improvement
Remote-flexible, offices in Toronto, New York, San Francisco, London and Paris, as well as a co-working stipend
6 weeks of vacation (30 working days!)
Company
Cohere
Cohere is an enterprise AI firm developing secure and private AI technology to address real-world business challenges.
H1B Sponsorship
Cohere has a track record of offering H1B sponsorships. Please note that this does not
guarantee sponsorship for this specific role. Below presents additional info for your
reference. (Data Powered by US Department of Labor)
Distribution of Different Job Fields Receiving Sponsorship
Represents job field similar to this job
Trends of Total Sponsorships
2025 (11)
2024 (14)
2023 (13)
2022 (5)
2021 (2)
Funding
Current Stage
Late StageTotal Funding
$1.71BKey Investors
Government of CanadaTiger Global ManagementIndex Ventures
2025-09-24Series D· $100M
2025-08-14Series D· $500M
2025-06-17Secondary Market
Recent News
2026-01-06
2025-12-27
2025-12-24
Company data provided by crunchbase