AI Benchmark Engineer | Native Language Specialist

AI Benchmark Engineer | Native Language Specialist

AI Benchmark Engineer | Native Language Specialist

Jobgether

2 horas atrás

Nenhuma candidatura

Sobre

  • This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI Benchmark Engineer | Native Language Specialist based in Brazil.
  • This is a remote freelance opportunity for a native-speaking software or prompt engineer passionate about multilingual AI and rigorous model evaluation. You will design realistic Terminal-Bench tasks that test how effectively large language models handle software challenges in your native language. Your work will help uncover failure modes related to multilingual prompts, non-English datasets, Unicode, encoding, localization, and terminal workflows. You will build and validate high-signal benchmark environments while ensuring that language-specific assets remain authentic rather than relying on English translations. Working across task engineering, implementation, calibration, and quality assurance, you will directly contribute to improving the reliability and multilingual robustness of AI systems. This role offers flexible, project-based work within a global community of technical and language experts.
  • n

Accountabilities

  • Design and engineer challenging, realistic benchmark tasks that evaluate coding agents and large language models in multilingual terminal environments.
  • Create authentic task environments using datasets, files, prompts, and other assets written in your native language, ensuring they accurately reflect real-world language usage.
  • Identify model failure points related to native-language prompting, translation gaps, multilingual reasoning, and language-specific software workflows.
  • Develop robust reference implementations and highly reliable, deterministic verifier scripts, using rubric-based evaluation only when strictly necessary.
  • Analyze execution logs and calibrate task difficulty from Easy to Very Hard using standardized Terminal-Bench configurations across different model tiers.
  • Participate in a rigorous multi-layer quality process covering task creation, human review, calibration review, and final audit, alongside automated LLM-based checks.
  • Validate grammatical accuracy, linguistic authenticity, technical correctness, fairness, reproducibility, and overall benchmark integrity.
  • Apply deep knowledge of multilingual text processing to identify edge cases involving Unicode normalization, encoding and decoding, locale behavior, text I/O, string operations, and toolchain interoperability.
  • Where relevant to the target language, account for bidirectional and RTL text handling, font fallbacks, rendering, and typography in software interfaces and generated artifacts.
  • Requirements
  • At least 1 year of professional experience in software engineering, prompt engineering, or a closely related technical field.
  • Demonstrated technical experience through work at established technology organizations and/or graduation from a strong engineering university.
  • Native or near-native fluency in the target language, with a sophisticated understanding of grammar, register, phrasing, and language-specific conventions.
  • Strong English proficiency for technical communication and collaboration.
  • Strong Python skills, along with proficiency in standard shell scripting and data processing workflows.
  • Extensive experience working with Terminal/CLI-based development environments and familiarity with coding agents or AI-assisted development tools.
  • Solid understanding of multilingual text-processing challenges, including encoding and decoding, Unicode normalization, locale-dependent casing and collation, non-Gregorian dates, text I/O, and safe string manipulation.
  • For applicable languages, familiarity with bidirectional or RTL text, font fallback behavior, and rendering or typography considerations.
  • Strong attention to detail and the ability to create deterministic, reproducible, technically rigorous evaluation tasks.
  • Ability to work independently, meet deadlines, and maintain consistent quality across project-based assignments.
  • Reliable availability and commitment once tasks are accepted, with most assignments requiring at least 2 hours per day or 10 hours per week.
  • Ability to provide an updated CV in English and successfully complete a GenAI assessment as part of the onboarding process.
  • Must be eligible to work as an independent contractor in the applicable location; geographic restrictions may apply in regions subject to international embargoes or sanctions.
  • As an independent contractor, you are responsible for your own applicable tax obligations.
  • Benefits
  • Fully remote, flexible freelance work that allows you to choose when and how much you work, without fixed hours or day-to-day micromanagement.
  • Competitive project-based compensation with prompt payments and a streamlined invoicing process.
  • Opportunity to contribute to cutting-edge AI evaluation and multilingual language technology with real-world impact.
  • Access to diverse technical and language-focused projects that can broaden your portfolio and strengthen your expertise across domains.
  • Collaboration with a global community of software engineers, linguists, language specialists, and AI experts.
  • Flexible supplemental work that can fit alongside other professional commitments, with project availability varying according to demand.
  • Exposure to emerging AI technologies, coding agents, benchmark development, and multilingual model evaluation.
  • No traditional employee benefits such as health insurance, paid time off, or retirement contributions, as this is an independent contractor opportunity.
  • No guaranteed hours or fixed workload; availability depends on project demand and successful task assignments.
  • n

How Jobgether works

  • We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
  • We appreciate your interest and wish you the best!
  • Why Apply Through Jobgether?
  • Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
  • #LI-CL1