Overview

Video Annotator – Indonesia Jobs in Indonesia at CNTXT AI

Title: Video Annotator – Indonesia

Company: CNTXT AI

Location: Indonesia

Video Annotator

Vision-Language-Action (VLA) Data Project 

Location: [Remote] 

Employment Type: [Full-time/ Contract] 

Experience Level: Entry to Mid-level 

About the Role 

We are looking for detail-oriented Video Annotators to support a Vision-Language-Action (VLA) dataset project. In this role, you will watch short video clips of physical tasks/activities and produce event-level annotations by segmenting each video into meaningful steps and writing clear, precise natural-language instructions describing the action taking place in each segment. 

Key Responsibilities 

  • Watch assigned videos in full before beginning annotation to understand the overall task and context.

  • Segment each video into discrete, logical steps/events based on visible action boundaries. 

  • Write clear, concise, and grammatically correct step-by-step descriptions (instructions) for each segment, accurately reflecting what is happening on screen. 

  • Ensure annotations are action-oriented, unambiguous, and consistent in tense, tone, and phrasing across the dataset. 

  • Assign accurate start and end timestamps for each segment. 

  • Follow annotation guidelines, taxonomies, and style guides provided by the project team.

  • Flag videos that are unclear, corrupted, mislabeled, or otherwise unsuitable for annotation. 

  • Participate in calibration sessions and incorporate reviewer/QA feedback to improve annotation quality and consistency. 

  • Meet daily/weekly productivity and quality targets without compromising accuracy.

  • Maintain confidentiality of all project data and materials. 

Required Qualifications 

  • English Proficiency: C1 level (CEFR) or above — required. Strong command of grammar, vocabulary, and sentence construction is essential, as output quality directly depends on writing clarity. 

  • Excellent attention to detail and ability to spot subtle changes in visual scenes. 

  • Strong written communication skills; ability to describe actions concisely and precisely.

  • Comfort working with video-based tools and annotation software (training provided).

  • Ability to work independently, follow detailed guidelines, and maintain consistency over repetitive tasks.

  • Reliable internet connection and access to a computer with a modern browser. 

  • Basic computer literacy (file handling, web-based tools, spreadsheets).

Preferred Qualifications 

  • Prior experience in data annotation, labeling, transcription, or content moderation. 

  • Familiarity with AI/ML concepts, especially computer vision, NLP, or robotics datasets.

  • Experience writing instructional or procedural content (e.g., how-to guides, SOPs, recipes).

  • Exposure to annotation platforms (e.g., CVAT, Label Studio, Scale, or similar tools). 

What We’re Looking For 

  • A sharp eye for detail and a passion for precision in written language. 

  • Someone who can watch an activity and translate it into clear, step-by-step written instructions a person (or a machine) could follow. 

  • A self-starter who is comfortable with repetitive, focused work and iterative feedback. 

Evaluation / Selection Process 

  1. English proficiency assessment (C1 level verification) 

  2. Annotation skills test (sample video segmentation + step-writing task) 

  3. Interview / calibration round 

  4. Onboarding and guideline training 

Compensation & Benefits 

$1 per accepted hours (Monday-Friday 9 hours a day) 

Upload your CV/resume or any other relevant file. Max. file size: 800 MB.