About the role
The Domain Scaling team works to make Claude world-class at real-world knowledge work in finance, healthcare, and legal. This role combines applied research and data sourcing — both real-world and synthetic — and owns the end-to-end creation of reinforcement learning (RL) environments: identifying high-value tasks, designing reward signals, managing data-vendor relationships, building QA frameworks to catch reward hacking, and measuring impact on model performance.
Responsibilities
- Own the end-to-end creation of RL environments for knowledge-work domains across finance, healthcare, and legal.
- Identify high-value, real-world tasks that are worth teaching a model to perform.
- Design reward signals that capture genuine task quality.
- Source data — both real-world and synthetic — and manage relationships with data vendors.
- Build QA frameworks to detect and prevent reward hacking.
- Measure the impact of environments on model performance and iterate.
Qualifications
- Strong software engineering skills, with proficiency in Python.
- Experience with machine learning, reinforcement learning, large language models, or building data and evaluation pipelines.
- Comfort working on ambiguous, open-ended research problems end to end.
- Interest in or familiarity with at least one applied domain: finance, healthcare, or legal.
- Experience designing evaluations, datasets, or reward and QA systems is a plus.
- Willing to work from the San Francisco or New York City office at least 25% of the time.
Responsibilities are taken from Anthropic's listing; qualifications are summarised for readability. For the authoritative, complete description and to apply, see Anthropic's posting via the Apply button above.
Apply
This is an external listing curated by InfoOnAIResources. Applications are handled entirely on Anthropic's own careers site: