You're seeing this page as if you were . The main menu is still yours, though. Exit from immersion
Jimmy BinghamJB

Jimmy Bingham

Senior AI Trainer and Annotator

€176/day
New York, US
3-7 years

Average response time: 1 hour

About Jimmy

AI systems don't train themselves , someone has to teach them right from wrong.

I'm an AI Trainer and LLM Evaluation Specialist with 5+ years of experience evaluating, red-teaming, and training large language models across Outlier, Scale AI, and Appen. I have a sharp eye for what makes AI responses accurate, safe, and genuinely useful. From ranking RLHF outputs to engineering adversarial prompts that expose model weaknesses — I sit at the intersection of human judgment and machine intelligence, making AI systems smarter one evaluation at a time.
  • English

    Native or bilingual

Remote only
Primarily works remotely

Experience

  • Outlier AI
    AI Trainer & Prompt Engineer
    January 2024 - Today (2 years and 7 months)
    • • Evaluated LLM outputs across coding, reasoning, math, and creative writing tasks, providing ranked comparisons and detailed written justifications to support RLHF training.
    • • Authored adversarial and edge-case prompts to stress-test model instruction-following, flagging failure modes that informed alignment improvements.
    • • Produced instruction-response pairs for supervised fine-tuning datasets, maintaining high consistency scores across quality audits.
    • • Collaborated with project leads to refine evaluation rubrics and calibration standards used across the annotator team.
  • Scale AI
    LLM Data Annotator
    March 2022 - December 2023 (1 year and 9 months)
    • • Completed thousands of model comparison tasks across NLP, summarization, and dialogue domains, ranking outputs by helpfulness, accuracy, and safety.
    • • Participated in red-teaming exercises designed to surface harmful, biased, or misleading model outputs for safety review teams.
    • • Maintained annotation accuracy scores consistently above platform quality thresholds across multiple concurrent projects.
    • • Provided detailed written feedback on model responses, helping researchers identify recurring error patterns and training gaps.
  • Appen
    Data Annotation Specialist
    June 2020 - February 2022 (1 year and 8 months)
    • • Annotated large volumes of text, audio, and image data for AI training projects across search relevance, sentiment analysis, and intent classification.
    • • Performed quality assurance reviews on fellow annotators' work, ensuring labeling guidelines were applied correctly and consistently.
    • • Contributed to multiple long-running AI data projects, demonstrating reliability and adaptability across shifting project requirements.
    • • Achieved top-tier contributor status based on accuracy metrics and throughput across consecutive project cycles.

Recommendations

Be the first to recommend Jimmy

Help this freelancer shine by sharing your experience working together.

These freelancer profiles also match your criteria

AgathaA

Agatha Frydrych

Backend Java Software Engineer

4.7

(3)

2

BaptisteB

Baptiste Duhen

Fullstack developer

4.6

(4)

5

AmedA

Amed Hamou

Senior Lead Developer

4

(2)

7

AudreyA

Audrey Champion

Web developer

4.3

(3)

4

Education

  • A.S. in Computer Science
    Los Angeles City College
    2017
    A.S. in Computer Science
  • Deep Learning Specialization – Coursera
    deeplearning.ai
    2022
    Deep Learning Specialization – Coursera

Categories