Founding Member of Technical Staff

San Francisco, CA · Full-time

About Arcophos

Arcophos builds the post-training layer for healthcare AI. We turn real clinical workflows into verifiable RL environments, with clinician-audited reward signals and episode-level evaluations, and use them to post-train models that enterprises and health systems can actually trust in rigorous, real-world care settings.

About the Role

As a Founding Member of Technical Staff at Arcophos, you'll build the systems that make healthcare models reliable: the post-training pipelines and the healthcare RL environments they learn from. You'll operate at the intersection of frontier ML research and clinical reality: turning real clinical workflows (EMR histories, triage and escalation pathways, ordering, documentation and coding) into verifiable training environments, then post-training models against them with reinforcement learning and preference optimization until they hold up over full patient episodes, not just single steps.

You'll move seamlessly between papers and production: leading large-scale experiments, building optimized training and evaluation pipelines, and setting the standard for what "clinically reliable" means in practice. As a founding member, your mandate goes beyond research. You'll shape the research agenda, the evaluation philosophy, the first hires, and the company itself, working directly with the founders with significant ownership and real compute behind you.

What You'll Do

  • Design and build healthcare RL environments from real clinical workflows, instrumenting intermediate milestones (escalation rules, timing targets, documentation requirements) to make traditionally unverifiable clinical tasks trainable
  • Author and audit the reward signals those environments emit: deterministic judges, LLM graders, and clinician-informed rubrics, with grader QA that catches false negatives and reward hacking before models learn to exploit them
  • Post-train healthcare models end to end: RL with verifiable rewards (GRPO-style climbs), reward modeling, preference optimization, synthetic data generation, and distillation
  • Build long-horizon evaluations that score dependent, multi-step clinical episodes, because per-step accuracy dramatically overstates real-world readiness
  • Stand up the supporting infrastructure: distributed training, high-throughput data pipelines, and large-scale experiment management
  • Work with clinicians and healthcare partners to source real workflow data and validate that environments reward what clinical experts actually value
  • Publish or contribute to leading-edge work in post-training and environment design
  • Help build Arcophos itself: research roadmap, recruiting, culture, and how we share our work

Requirements

  • Deep experience in machine learning, ideally reinforcement learning, post-training, or alignment research
  • Demonstrated research contributions: published papers (NeurIPS, ICML, ICLR) or strong public implementations
  • Strong proficiency in Python and ML frameworks (PyTorch or JAX)
  • Comfort with distributed training, high-throughput data pipelines, and large-scale experiment management
  • Ability to reason independently and run the full loop: hypothesis → experiment → insight → shipped improvement
  • Founding-member disposition: high agency, comfort with ambiguity, and willingness to build infrastructure one week and read clinical guidelines the next

Nice to have

  • Experience with healthcare data (EMR/EHR, FHIR, claims) and its privacy constraints (HIPAA, de-identification)
  • Prior work designing RL environments, evals, graders, or reward models
  • Experience collaborating with clinicians or other domain experts

Compensation & Benefits

  • Base: $200,000-$350,000+
  • Significant founding equity
  • Full medical, dental, and vision coverage
  • Free daily meals (breakfast, lunch, and dinner)
  • Free gym membership
  • Unlimited book budget (help us build our library)
  • $10,000+ housing stipend

Apply for this role