at Handshake
Handshake is seeking a remote US-based AI Policy Generalist to turn complex customer policies into consistent, well-reasoned evaluations of AI model behavior.
ABOUT HANDSHAKE
Handshake was founded on a simple belief that everyone deserves a path to a great career, regardless of where they went to school or who they know. Today, we power 25 million job seekers, 1 million+ employers, and 1,600 educational institutions.
In 2025, we started Handshake AI and built the fastest-growing AI data business in history. We work directly with frontier AI lab researchers to create evaluations, publish benchmarks, and push the boundary of data. We’ve grown from $0 to ~$1B run rate and pay ~$60M to over 30K individuals every month.
Why join Handshake now
About Handshake AI
Human data is the core infrastructure to AI advancement. Frontier AI labs currently improve model capabilities with various data-intensive post-training techniques. We believe that data spend for AI training will increase by 3-5x in the next few years and continue for much longer as models take on new domains. Handshake AI supports all of the frontier AI labs, working on their most complex data at the largest scale.
ABOUT THE ROLE
As an AI Policy Generalist, you will turn complex customer policies into consistent, well-reasoned evaluations of AI model behavior.
You will read user requests, model responses, and relevant conversation history, then determine which policy category best applies. The most interesting cases will not have obvious answers. Two examples may look almost identical until a single word, contextual detail, or difference in intent changes the correct classification.
We are looking for people who enjoy splitting hairs in a healthy way. You form clear opinions, explain precisely why two cases should be treated differently, challenge interpretations respectfully, and change your mind when better evidence emerges. You understand that productive disagreement is not about winning an argument. It is how a team finds the most accurate and consistent interpretation.
This is not rote annotation. Policies cannot anticipate every possible edge case, and good evaluators do not apply them mechanically. You will balance the policy’s text and intent with customer expectations, conversation context, precedent, and team calibration.
The subject matter will vary. One project may involve distinguishing benign assistance from meaningful facilitation of harm. Another may require evaluating whether an interaction reflects ordinary emotional support or unhealthy reliance. A third may focus on nuanced boundaries within sexual-safety policy. Success requires learning each customer’s framework on its own terms rather than carrying assumptions from one domain into another.
WHAT YOU WILL DO
Responsibilities