Role Overview: We are seeking Bioinformatics and Computational Genomics experts to design challenging, agentic genomics tasks and reference solutions for GeneBench Pro. You will create realistic computational biology problems that require an AI agent to interpret biological data, write and execute code, work across scientific files, and produce objectively verifiable outputs. Tasks should reflect authentic genomics and bioinformatics workflows rather than standalone biology questions. Environment: Each GeneBench Pro problem is a self-contained scientific analysis. Agents receive an isolated workspace with a short prompt, data files, and a standard bioinformatics stack including Python, scientific computing libraries, and basic genomics packages such as PLINK 2.0. Reference solutions must run in this Python-based environment. Key Responsibilities: Design novel, model-challenging tasks in computational genomics and bioinformatics. Design tasks that test higher-order scientific judgment, including handling ambiguous or messy data, identifying artifacts, selecting appropriate analytical approaches, revising assumptions based on intermediate results, and determining when conclusions are decision-ready. Build tasks involving realistic scientific data such as FASTA/FASTQ, VCF, BAM/SAM, BED, TSV/CSV, sequence annotations, expression data, germline or somatic variant data, or other genomics artifacts. Develop tasks requiring multi-step analysis using Python, command-line tools, or established bioinformatics libraries. Create clear task specifications, input datasets, expected output schemas, and deterministic or objectively verifiable ground truths. Develop expert reference solutions and reproducible computational workflows that run in the provided Python stack. Validate that tasks are scientifically correct, solvable from the supplied information, and sufficiently challenging for frontier AI models. Design robust grading criteria that distinguish scientifically correct solutions from superficially plausible outputs. Ensure all deliverables are well documented, reproducible, and client-ready. Maintain high quality and throughput while incorporating reviewer feedback. Communicate progress, blockers, and scientific or technical requirements to project leads and reviewers. Qualifications Required: Ph.D., postdoctoral experience, or equivalent research experience in Bioinformatics, Computational Biology, Genomics, Computational Genetics, or a closely related discipline. Strong hands-on programming experience in Python. Experience analyzing biological sequence or genomics datasets. Comfortable working in Linux/command-line computational environments. Preferred: Experience with genomics workflows such as variant analysis, transcriptomics, sequence analysis, phylogenetics, population genetics, functional genomics, or clinical genomics. Familiarity with common bioinformatics libraries and tools such as Biopython, pandas, NumPy, SciPy, samtools, bcftools, BLAST, PLINK, or equivalent tools. Experience developing reproducible scientific pipelines. Experience evaluating AI/LLM systems on computational scientific tasks. Strong understanding of experimental and biological context behind computational analyses. Bonus Points For Experience with AI agents or coding agents. Experience designing benchmark datasets or automated graders. Publications involving computational genomics or bioinformatics. Experience with Docker/containerized scientific workflows. Offer Details: Remote Full-time dedication: 40 hours per week with 4 hours PST overlap per day Location: Argentina, Bangladesh, Brazil, Colombia, Egypt, Ghana, India, Indonesia, Turkey, Chile, Peru, Sri Lanka, Saudi Arabia, Nepal, Mexico Engagement Length: 5 weeks Please note: This position is paid on a task based model. The hourly rate listed in this form is only an estimated average to provide compensation context. The final payment will depend on the tasks completed. Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/30/2026
Role Overview We are seeking Physical Sciences experts to develop realistic, terminal-based scientific tasks for Terminal Bench Science. You will translate authentic physics, chemistry, materials science, astronomy, and computational science workflows into reproducible benchmark environments. The role involves creating scientific inputs, computational models, executable solutions, automated tests, and objective grading criteria. You will help evaluate whether AI agents can reason through scientific problems, operate command-line tools, debug calculations, and produce reliable scientific artifacts. Core Domains Physics, chemistry, materials science, computational physics, computational chemistry, astronomy and cosmology, thermodynamics, quantum mechanics, statistical mechanics, or materials modeling. What you'll do Design multi-step terminal tasks based on realistic physical-science workflows. Develop self-contained computational environments with pinned dependencies and scientific software. Create input datasets, molecular structures, simulation parameters, experimental data, or model configurations. Implement expert solutions using Python, Bash, C/C++, Julia, or domain-specific tools. Develop automated tests for numerical accuracy, physical consistency, convergence, and output structure. Create tasks involving simulations, numerical modeling, data fitting, optimization, spectroscopy, molecular analysis, or scientific visualization. Define appropriate numerical tolerances, units, boundary conditions, and expected scientific behavior. Validate that tasks are reproducible and execute successfully without runtime downloads. Debug issues involving dependencies, precision, solver stability, performance, and file formats. Communicate scientific assumptions and computational limitations clearly to reviewers. What we're looking for Ph.D., postdoctoral experience, or equivalent advanced technical experience in a relevant physical-science discipline. Strong programming skills in Python, C/C++, Julia, Bash, or another scientific programming language. Experience working in Linux or terminal-based environments. Experience with numerical methods, scientific modeling, simulations, or quantitative data analysis. Ability to create and validate computational scientific workflows independently. Strong understanding of units, numerical precision, physical constraints, and scientific reproducibility. Nice to have Experience with NumPy, SciPy, pandas, matplotlib, SymPy, JAX, or similar libraries. Familiarity with molecular dynamics, quantum chemistry, finite-difference methods, Monte Carlo methods, optimization, or statistical modeling. Experience with tools such as OpenMM, ASE, RDKit, Psi4, LAMMPS, GROMACS, or other domain-specific software. Familiarity with Docker, Conda, Git, CI systems, and automated testing. Experience working with HPC systems or performance-sensitive scientific workloads. Bonus Points Experience evaluating AI coding or terminal agents. Research software engineering experience. Experience creating benchmark tasks or automated graders. Publications or open-source contributions involving computational science. Experience converting experimental or research workflows into reproducible packages. Offer Details: Remote Full-time dedication: 40 hours per week 4 hours PST overlap per day Location: Bangladesh, India, Indonesia, Egypt, Ghana, Turkey, Brazil, Colombia Engagement Length: 5 weeks Please note: This position is paid on a task based model. The hourly rate listed in this form is only an estimated average to provide compensation context. The final payment will depend on the tasks completed. Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/30/2026
Role Overview: We are seeking Mathematics experts to develop realistic, terminal-based scientific tasks for Terminal Bench Science. You will design tasks involving numerical analysis, optimization, statistics, mathematical modeling, probability, dynamical systems, and computational mathematics. The role focuses on evaluating whether AI agents can formulate mathematical problems, implement reliable algorithms, operate in a terminal environment, debug numerical workflows, and produce verifiable computational results. Core Domains: Applied mathematics, numerical analysis, optimization, statistics, probability, mathematical modeling, differential equations, dynamical systems, operations research, and computational geometry. What you'll do: Design authentic, multi-step computational mathematics tasks. Translate mathematical and research workflows into self-contained terminal environments. Prepare datasets, equations, model definitions, constraints, initial conditions, and expected outputs. Implement expert solutions using Python, R, Julia, C/C++, Bash, or other relevant tools. Create tasks involving optimization, numerical integration, differential equations, matrix computation, statistical inference, stochastic modeling, or algorithm analysis. Define rigorous grading criteria based on numerical accuracy, convergence, complexity, feasibility, and mathematical correctness. Establish appropriate tolerances, stopping criteria, stability requirements, and reproducibility controls. Develop automated tests that validate results across edge cases and alternative valid implementations. Debug issues involving floating-point precision, solver failures, conditioning, convergence, and performance. Document assumptions, mathematical formulations, expected outputs, and known limitations. What we're looking for: Ph.D., postdoctoral experience, or equivalent advanced technical experience in mathematics, statistics, or a closely related discipline. Strong programming skills in Python, R, Julia, C/C++, Bash, or another relevant language. Experience working in Linux or terminal-based environments. Experience with numerical methods, optimization, statistics, mathematical modeling, or scientific computation. Ability to independently implement, test, and validate computational algorithms. Strong understanding of numerical stability, error analysis, mathematical assumptions, and reproducibility. Nice to have Experience with NumPy, SciPy, SymPy, pandas, scikit-learn, JAX, PyTorch, CVXPY, or similar tools. Familiarity with optimization solvers, probabilistic programming, differential-equation libraries, or numerical linear algebra packages. Experience with Monte Carlo methods, stochastic processes, Bayesian inference, graph algorithms, or computational geometry. Familiarity with Docker, Conda, Git, CI systems, and automated testing. Experience developing mathematical programming challenges, benchmark tasks, or automated graders. Bonus Points Experience evaluating AI coding or terminal agents. Research software engineering experience. Experience working with HPC or performance-critical numerical code. Publications or open-source contributions in applied mathematics or computational science. Experience designing tests that distinguish mathematically valid solutions from approximate or numerically unstable ones. Offer Details: Remote Full-time dedication: 40 hours per week with 4 hours PST overlap per day Location: Bangladesh, India, Indonesia, Egypt, Ghana, Turkey, Brazil, Colombia Engagement Length: 5 weeks Please note: This position is paid on a task based model. The hourly rate listed in this form is only an estimated average to provide compensation context. The final payment will depend on the tasks completed. Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/30/2026
We’re looking for a Product Designer to help shape experiences across a large-scale digital advertising ecosystem. You’ll focus on areas such as campaign management, reporting, account management, billing, permissions, and identity, translating complex data and workflows into clear, intuitive experiences. You’ll design for both web and mobile, helping advertisers of varying sizes and levels of sophistication understand performance, optimize outcomes, and manage their accounts effectively. Working closely with product managers, engineers, researchers, data scientists, and other cross-functional partners, you’ll contribute throughout the full design lifecycle—from early exploration and systems thinking to detailed workflows and polished implementation. You’ll also help identify strategic opportunities, simplify complex workflows, build user trust, and drive alignment through thoughtful storytelling, communication, and presentations. Required Skills 4+ years of professional UX/UI or Product Design experience. Strong portfolio demonstrating product design work across web and mobile applications. Advanced Figma skills. Significant hands-on experience using AI tools such as Claude and/or Cursor as part of the design workflow. Strong product sensibilities and experience designing intuitive solutions for complex workflows. Experience contributing across the full product design lifecycle, from exploration and systems thinking through final delivery. Ability to navigate ambiguity and bring clarity to complex product and data challenges. Strong communication, storytelling, and presentation skills. Experience collaborating closely with product managers, engineers, researchers, data scientists, and other cross-functional stakeholders. Ability to design effectively for users with varying levels of technical sophistication. Bonus Skills Experience across both B2B and B2C products. Experience designing advertising platforms, campaign management tools, reporting products, or similar advertiser-facing experiences. Familiarity with account management, billing, permissions, identity, or data-heavy product experiences. Offer Details Remote Full-time dedication (40 hours/week) Location: USA Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/30/2026
Role Overview: As an AI Quality Analyst, you will evaluate a new personalization feature for Gemini. You will assess how well the model uses information from your past Gemini conversations, Gmail, Google Search, and YouTube activity to make responses more relevant and helpful. This role requires a unique blend of creativity and analytical rigor. You will actively design prompts from the perspective of your own personal experiences. You will then use your analytical skills to assess the quality of the model's personalized responses, evaluating dimensions like Grounding, Integration, and Helpfulness. Key Qualifications: English Proficiency: Ability to read and write in English with a high degree of comp, as English is the focus language for this project. Personal Account Usage: Willingness to use your primary personal Google account (not a testing account) and enable personal data sources for a genuine assessment. Schedule Flexibility: Full-time availability in your local time zone is required. We are staffing a global, 24-hour operations team. Exceptional Analytical Thinking: Demonstrate ability to evaluate nuanced and ambiguous AI responses, specifically assessing personalization quality. Creative Prompt Engineering: Experience in designing creative, multi-turn starting prompts based on personal context to thoroughly test the model's capabilities. Strong Evaluation Acumen: Understanding of personalization concepts, including the ability to identify incorrect personalization, poor inferences, and forced connections. Meticulous Attention to Detail: The ability to review Side-by-Side (SxS) model responses and spot subtle differences in naturalness and overnarrating. Excellent Written Communication: Superior ability to write clear, concise, and structured rationales for model rankings, explicitly referencing specific turn numbers. Feedback: Ability to provide constructive feedback and detailed annotations. Communication: Excellent communication and collaboration skills. Independence: Self-motivated and able to work independently in a remote setting. Technical Setup: Desktop/Laptop set up with a good internet connection. Description: In this role, you will be part of a dynamic team focused on evaluating the quality of personalized AI interactions. Your day-to-day work will involve: Designing and executing multi-turn conversational prompts (typically 1-5 turns) that require the AI to utilize your personal information and experiences. Evaluating model responses based on your intent from the starting prompt, checking if the personalization was appropriately applied. Analyzing responses for Grounding issues, ensuring claims about you are supported by evidence and not flawed inferences or hallucinations. Assessing Integration quality to ensure personal data is woven naturally into the response without robotic "overnarrating". Rigorously evaluating and stack-ranking two model responses side-by-side (SxS) to determine which is overall more helpful, easy to use, and enjoyable. Writing clear, defensible rationales for your comparisons, explicitly referencing where issues or positive aspects occurred in the conversation. Extracting and verifying "Debug Info" from the model to confirm that chat summaries and data sources were properly utilized. Maintaining strict data hygiene by deleting evaluation conversations to prevent them from polluting your future chat history. Education & Experience BS/BA degree or equivalent experience in a relevant field (e.g., Policy, Law, Ethics, Linguistics, Journalism, Computer Science, or a related analytical field). Experience in data annotation, AI quality evaluation, content moderation, or a related role is strongly preferred. Offer Details: Remote Full Time Availability (8 Hrs) and 4 hours overlap with the PST time zone for US Location: United States Engagement Type: Short Term Contract (4 month) Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/29/2026
Role Overview: We’re seeking creative and charismatic storytellers to join a cutting-edge initiative designed to train AI systems to better understand and describe the world in a natural, engaging way. As a Video Creator, you’ll act as a local narrator, producing short, dynamic “walk-and-talk” videos that showcase interesting locations near you. You’ll help capture authentic, human perspectives that reflect the character, diversity, and beauty of your surroundings. What You’ll Do Day-to-Day: Record 1–5 minute first-person "walk-and-talk" videos in real-world environments while completing a single task or mission. Interact naturally with an AI assistant by giving short voice commands as you point to or handle objects, following the recording guidelines. Record smooth, steady footage using your smartphone or camera with clear audio. Choose safe, public locations where filming is legally allowed. Follow provided creative guidelines, ensuring each video has a clear introduction, middle section, and conclusion. Submit videos along with accurate location details (name and address or coordinates). Respect privacy and safety best practices when filming in public areas. Requirements: Comfortable speaking on camera in English with clear communication. Creative, personable, and confident presence on video. Owns a smartphone or recording device with good video and audio quality. Reliable internet connection for uploading videos. Ability to follow task guidelines and deliver videos on schedule. Nice to Have: Background or interest in vlogging, photography, tourism, or performing arts. Experience creating videos for YouTube, TikTok, or Instagram Reels. Familiarity with basic video editing or framing techniques. Offer Details: Remote Commitments Required: At least 4 hours per day and minimum 20 hours per week with overlap of 4 hours with PST. Location: United States Engagement Length: 3 months Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/28/2026
About the Role: We're looking for a high-taste engineer with a strong eye for quality — someone who's worked hands-on with preference data and rubric creation, and knows what separates a good example from a great one. You communicate clearly and can synthesize complex technical work the way a strong MAI dev would: concise, structured, and easy for others to act on. You're comfortable in ambiguous, fast-moving environments and default to shipping over deliberating. What does day-to-day look like: Design, refine, and iterate on rubrics and evaluation criteria for preference data Review and label data with a high bar for quality, catching subtle issues others might miss Build and maintain pipelines/infra that support data generation, collection, and evaluation workflows Synthesize findings from data work into clear write-ups, updates, or recommendations for the team Collaborate closely with researchers/engineers to translate qualitative judgment into scalable processes Required Skills: 5+ years of proficiency in Python (especially building/maintaining pipelines and infra) or Typescript/Golang/JavaScript (experience contributing to or navigating large codebases like VSCode a plus) Hands-on experience with preference data collection and/or rubric creation Strong written and verbal communication — able to distill technical work into clear, actionable summaries Demonstrated high-taste judgment — able to articulate why something is good, not just that it is Offer Details: Remote Commitments Required: At least 4 hours per day and minimum 20 hours per week with overlap of 4 hours with PST. Location: United States Engagement Type: Short Term Contract (3 month) Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/28/2026
This is a long-term engagement with a US-based client in the employee benefits technology space. Openings on this requisition: 10 Location: Remote, anywhere in Brazil Engagement: Long-term project with a US-based client About the role As a Software Engineer, you will work in a team that partners closely with Product Managers and customers to learn about the challenges employers face while navigating the competitive employee benefits landscape. You will design solutions that solve problems in ways our customers love and work for our business. You will build the highest quality software in the latest technologies and test driven development practices. Responsibilities Collaborate with stakeholders to learn about our customers biggest challenges. Measure, inspect, and drive decisions using data. Design, test, code, and instrument new solutions, with guidance and code review from senior engineers. Follow the team's engineering process: TDD and BDD, Microservice and Vertical Slice Architectures. Support live applications, take part in proactive monitoring, incident response, and continuous improvement. Learn from your peers and take part in the continuous learning of your team. Grow your knowledge of your functional area and its best practices. Apply problem-solving techniques to resolve issues and suggest approaches. Complete well-defined tasks and proactively review your work with others. Self-motivated, take ownership of your work, and actively seek out ways to contribute. Required Qualifications Bachelor's degree in Computer Science, Software Engineering, or related field, OR demonstrable equivalent experience. 0 to 2 years of experience in software engineering (internships, academic and personal projects count). Strong problem-solving and analytical skills. Good communication and collaboration skills. Passionate about keeping up with modern technologies and design. Familiarity with a modern web UI framework (Angular or React), or willingness to learn one quickly. Basic understanding of REST APIs and how to build and consume them. Willingness to write unit tests and learn test driven development. Awareness of software security basics (OWASP guidelines). Familiarity with Git version control. Familiarity with Agile development methodologies. Basic SQL: queries, joins, and the fundamentals of a relational database. Technology Must-Haves C# Modern RDBMS (MS SQL, Postgres, MySQL) ASP.NET RESTful API basics Technology Nice-To-Have or Dedicate to Learning Quickly Docker Event-driven design with Kafka or an equivalent messaging platform (Azure Service Bus, RabbitMQ, SQS) Modern Web UI frameworks and libraries (Angular, React) Kubernetes Helm / ArgoCD Terraform GitHub Actions (the CI/CD pipeline you will work with) NoSQL databases GraphQL Azure Cloud-Native applications and services
Posted 9/24/2026
Role Overview We are seeking experienced legal professionals to generate, evaluate, and refine expert-level tasks related to the work of lawyers. The role requires strong legal expertise and the ability to produce technically accurate, practical, and high-quality legal tasks and solutions. Role Description: Create expert-level legal tasks and solutions across contract, litigation-support, and legal-research workflows. Evaluate and refine tasks to meet the pod’s quality standards for accuracy, clarity, and practical applicability. Work under senior review and incorporate calibration feedback to continuously improve task quality. Apply rubric standards and guidance established by the Functional SME (Expert) Lead. Deliver consistent, high-quality legal outputs while meeting task-specific requirements and deadlines. Preferred Qualifications: 3–8 years of legal practice, preferably at the Associate level. JD or equivalent legal qualification with active US bar admission. Strong experience in contract drafting, litigation support, and legal research. Ability to analyze complex legal issues and translate them into practical, high-quality tasks. Strong attention to detail, written communication, legal reasoning, and ability to incorporate feedback. Offer Details: Remote Full-time dedication - 40 hours per week with 8 hours PST overlap Location: United States Engagement Length: 6 weeks Please note: This position is paid on a task based model. The hourly rate listed in this form is only an estimated average to provide compensation context. The final payment will depend on the tasks completed. Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/23/2026
About the role We are seeking Life Sciences Subject Matter Experts to develop advanced scientific reasoning tasks for LifeSciBench. Unlike coding-focused benchmarks, this role emphasizes deep biological knowledge, scientific interpretation, and rigorous reasoning. You will create challenging questions based on realistic scientific scenarios, experimental results, biological data, figures, literature-style evidence, and research problems. This is a non-code benchmark: tasks are not executed, and no programming is required. Who this is for: Life scientists who want to author rigorous scientific-reasoning problems with objectively defensible answers — no coding involved. If your strength is interpreting experiments, spotting ambiguity, and constructing questions that reward synthesis over recall, this is the fit. Who this isn't for: If you want to build tasks where agents write and execute code, see GeneBench Pro (genomics), Agentic Life Sciences (broad life-sciences workflows), or Terminal Bench Science (cross-STEM terminal tasks). Key Responsibilities Design novel, challenging life sciences problems requiring advanced scientific reasoning. Create tasks grounded in realistic research scenarios, experiments, biological datasets, figures, sequences, structures, or scientific evidence. Develop clear questions with well-defined and objectively defensible answers. Write rigorous expert reference answers explaining the scientific reasoning required. Ensure tasks require synthesis and interpretation rather than simple factual recall. Validate all biological claims, calculations, assumptions, and conclusions. Identify ambiguity, alternative scientifically valid interpretations, and insufficiently constrained questions. Create tasks capable of distinguishing frontier AI models based on scientific reasoning ability. Ensure outputs are clearly documented and reproducible where applicable. Maintain high quality and throughput while incorporating reviewer feedback. Note on inputs: Tasks may involve scientific text, structured data, sequences, experimental results, and other scientific evidence depending on project requirements. Requirements Ph.D., postdoctoral experience, or equivalent advanced research experience in the Life Sciences. Strong expertise in at least one biological discipline. Demonstrated ability to interpret scientific experiments, data, and research findings. Relevant backgrounds include: Molecular Biology Genetics and Genomics Cell Biology Biochemistry Neuroscience Microbiology Immunology Evolutionary Biology Ecology Structural Biology Biophysics Developmental Biology Cancer Biology Pharmacology or related disciplines Preferred Qualifications Experience designing graduate-level scientific problems or assessments. Strong familiarity with experimental design and interpretation. Experience reviewing scientific manuscripts or research. Ability to translate complex scientific concepts into precise, self-contained problems. Experience evaluating LLM-generated scientific responses. Bonus Points For Publications in strong peer-reviewed journals or conferences. Teaching or mentoring experience at the graduate level. Experience with AI/ML evaluation datasets. Offer Details: Remote 40 hours per week with at least 4 hours PST overlap Location: Argentina, Bangladesh, Brazil, Colombia, Egypt, Ghana, India, Indonesia, Turkey, Chile, Peru, Sri Lanka, Nepal, Mexico Engagement Length: 4 weeks Please note: This position is paid on a task based model. The hourly rate listed in this form is only an estimated average to provide compensation context. The final payment will depend on the tasks completed. Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/22/2026
Responsibilities: Validate task quality: Check that instructions, source materials, reference solutions, and evaluation criteria are consistent, with no hidden requirements or missing information. Review agent performance: Inspect execution traces, tool calls, and generated deliverables to determine whether successes and failures are justified. Audit grading logic: Identify brittle checks, incorrect expected answers, unsupported rubric criteria, and cases where valid alternative solutions are unfairly penalized. Investigate discrepancies: Distinguish genuine model limitations from task defects, grader errors, and environment or tool failures. Assess automated QC findings independently rather than accepting them at face value. Document decisions: Provide concise, evidence-backed findings and actionable feedback, flag uncertainty, and verify that revisions resolve identified issues. What we're looking for: Technical fluency: Comfortable reading Python, SQL, shell scripts, structured data, and execution logs to understand task setup and grading behavior. Analytical judgment: Able to independently check calculations, reconcile conflicting evidence, and assess the correctness and completeness of professional deliverables. Clear communication: Strong written English, attention to detail, and experience providing specific, reproducible feedback. Relevant experience preferred: AI evaluation, technical QA, data analysis, or benchmark development; familiarity with Harbor task setup. Offer Details: Remote Full-time dedication - 40 hours per week with 8 hours PST overlap per day Location: LATAM Engagement Length: 10 weeks Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/22/2026
Role Overview We are seeking Computational Life Sciences experts with strong scientific programming skills to develop complex, realistic tasks for Agentic Life Sciences. You will design tasks where AI agents must independently navigate scientific data, write and execute code, use computational tools, troubleshoot intermediate results, and produce scientifically valid final outputs. The goal is to evaluate whether AI systems can perform authentic multi-step scientific work — not simply answer scientific questions. Environment Tasks are executed in controlled computational environments and should be self-contained and reproducible. Depending on the scientific workflow, tasks may use Python, command-line tools, scientific libraries, or domain-specific software. Experts should design tasks whose required dependencies, data, and computational resources can be reliably packaged and evaluated in the project environment. Key Responsibilities Design challenging, realistic agentic scientific workflows across the life sciences. Create tasks requiring multiple computational steps rather than single-script or simple question-answer solutions. Develop realistic input files, scientific datasets, instructions, constraints, and expected deliverables. Build tasks requiring agents to inspect data, select appropriate methods, execute analyses, troubleshoot problems, and synthesize results. Create reproducible expert solutions and objectively verifiable ground truths that run in the sandbox described above. Design robust automated or semi-automated grading criteria. Ensure tasks measure scientific reasoning and execution rather than memorization. Validate scientific assumptions, calculations, code, intermediate outputs, and final answers. Ensure tasks are self-contained and executable in controlled, network-isolated computational environments. Maintain high quality and throughput while responding effectively to reviewer feedback. Qualifications Required Ph.D., postdoctoral, or equivalent research experience in the life sciences, with strong demonstrated computational and scientific-programming experience. Strong scientific programming experience, particularly in Python. Experience performing multi-step computational scientific analyses. Ability to independently validate both scientific reasoning and computational outputs. Preferred Expertise in one or more areas including bioinformatics, computational genomics, systems biology, computational neuroscience, biostatistics, computational drug discovery, computational biochemistry, structural biology, protein engineering, or computational microbiology. Experience with scientific libraries and command-line tools. Experience creating reproducible research pipelines. Ability to translate authentic research workflows into bounded, objectively gradable tasks. Bonus Points For Experience with AI agents, coding agents, or scientific AI systems. Experience building automated evaluation environments. Familiarity with Docker/Linux environments. Publications involving computational or data-intensive life sciences research. Offer Details: Remote Full-time dedication - 40 hours per week with 4 hours PST overlap per day Location: Argentina, Bangladesh, Brazil, Colombia, Egypt, Ghana, India, Indonesia, Turkey, Chile, Peru, Sri Lanka, Saudi Arabia, Nepal, Mexico Engagement Length: 5 weeks Please note: This position is paid on a task based model. The hourly rate listed in this form is only an estimated average to provide compensation context. The final payment will depend on the tasks completed. Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/22/2026
Role Overview We are seeking Earth Sciences experts to develop realistic, terminal-based scientific tasks for Terminal Bench Science. You will create computational workflows involving environmental data, climate systems, atmospheric processes, geophysics, oceanography, geology, and related fields. You will design tasks that require AI agents to inspect scientific datasets, process geospatial or time-series data, run models, troubleshoot pipelines, and generate objectively verifiable scientific outputs. What you'll do Translate authentic Earth-science workflows into self-contained terminal benchmark tasks. Prepare geospatial, climate, atmospheric, geological, hydrological, or oceanographic datasets. Build reproducible computational environments with appropriate scientific libraries and command-line tools. Create expert solutions using Python, R, Bash, Julia, or domain-specific software. Develop tasks involving geospatial analysis, time-series processing, numerical modeling, interpolation, forecasting, remote sensing, or environmental risk analysis. Define objective grading criteria for scientific outputs, model behavior, data transformations, and spatial or temporal accuracy. Validate coordinate systems, units, timestamps, missing-data handling, and scientific assumptions. Create automated tests for numerical tolerances, file formats, metadata, and reproducibility. Debug issues involving geospatial projections, large datasets, dependencies, performance, and numerical stability. Document input data provenance, expected outputs, edge cases, and limitations. What we're looking for Ph.D., postdoctoral experience, or equivalent advanced technical experience in Earth Sciences or a closely related field. Strong expertise in at least one area such as climate science, atmospheric science, geophysics, oceanography, geology, hydrology, remote sensing, environmental modeling, or Earth-system science. Strong programming skills in Python, R, Julia, Bash, or another scientific programming language. Hands-on experience with scientific data processing, numerical modeling, geospatial analysis, environmental datasets, or time-series analysis. Comfortable working independently in Linux/terminal-based environments. Ability to build, debug, and validate reproducible scientific computational workflows. Strong understanding of scientific quality control, spatial/temporal data, uncertainty, and numerical accuracy. Preferred Qualification Experience with NumPy, pandas, SciPy, xarray, rasterio, GeoPandas, Cartopy, GDAL, or similar scientific/geospatial tools. Experience with scientific data formats such as NetCDF, HDF5, GeoTIFF, shapefiles, or GRIB. Experience working with climate, weather, satellite, seismic, oceanographic, geological, or hydrological datasets. Familiarity with Docker, Conda, Git, CI/CD, automated testing, or HPC environments. Nice to Have Research software engineering, scientific benchmarking, or automated grader development experience. Experience evaluating AI coding/terminal agents or developing tasks and evaluations for AI systems. Publications or open-source contributions in Earth, environmental, geospatial, or computational sciences. Offer Details Remote Full-time dedication - 40 hours per week with 4 hours PST overlap per day Location: Bangladesh, India, Indonesia, Egypt, Ghana, Turkey, Brazil, Colombia Engagement Length: 5 weeks Please note: This position is paid on a task based model. The hourly rate listed in this form is only an estimated average to provide compensation context. The final payment will depend on the tasks completed. Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/22/2026
Role Overview We are staffing a frontier AI initiative that requires strong Engineering experts to develop realistic, terminal-based scientific and technical tasks used to train and evaluate AI agents. The work involves translating authentic engineering workflows into reproducible computational tasks that test modeling, simulation, optimization, data processing, debugging, and technical validation. The role spans disciplines including mechanical, electrical, chemical, aerospace, civil, materials, biomedical, robotics, and control systems engineering. What you'll do: Design realistic, multi-step terminal tasks based on real-world engineering and scientific workflows. Create engineering datasets, simulation inputs, geometry files, sensor data, design constraints, and configuration files. Develop expert solutions using Python, C/C++, Julia, MATLAB/Octave, Bash, or relevant engineering software. Build reproducible, containerized environments with appropriate engineering tools and pinned dependencies. Develop tasks involving simulation, numerical analysis, optimization, control systems, signal processing, finite-element concepts, CAD-related data, and engineering design. Create automated tests and objective grading criteria that validate engineering correctness, including units, physical constraints, tolerances, convergence, stability, boundary conditions, and numerical behavior. Debug solver, dependency, workflow, precision, and performance issues and clearly document assumptions, requirements, expected outputs, and edge cases. What we're looking for: Ph.D., postdoctoral experience, or equivalent advanced technical experience in an Engineering discipline such as mechanical, electrical, chemical, aerospace, civil, materials, biomedical, robotics, or control systems. Strong scientific programming skills in Python, C/C++, Julia, MATLAB/Octave, Bash, or another relevant language. Hands-on experience working in Linux or terminal-based environments. Experience with engineering simulation, modeling, numerical analysis, optimization, signal processing, control systems, or technical data analysis. Strong understanding of numerical methods, engineering units, physical constraints, boundary conditions, and technical validation. Ability to build, debug, and validate reproducible computational engineering workflows. Nice to have Experience with scientific libraries such as NumPy, SciPy, pandas, matplotlib, SymPy, or PyTorch. Familiarity with engineering tools such as OpenFOAM, CalculiX, FEniCS, ROS, LTspice-compatible workflows, QEMU, or similar software. Experience with finite-element methods (FEM), computational fluid dynamics (CFD), robotics, embedded systems, control systems, or digital twins. Familiarity with Docker, Git, CI/CD, automated testing, or HPC environments. Experience developing technical benchmarks, programming tasks, simulation-based evaluations, or automated graders. Experience evaluating AI coding or terminal agents or working in research software engineering. Publications, patents, open-source contributions, or industry experience involving computational engineering. Offer Details: Remote Employment type: Contractor assignment Full-time dedication - 40 hours/week with 4 hours PST overlap per day Location: Bangladesh, India, Indonesia, Egypt, Ghana, Turkey, Brazil, Colombia Engagement Length: 5 weeks Please note: This position is paid on a task based model. The hourly rate listed in this form is only an estimated average to provide compensation context. The final payment will depend on the tasks completed. Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/22/2026
Right hand to the Founder & CEO | Remote from Brazil | US startup Mavila Consulting is hiring for a fast-growing US startup that has grown double digits every single year. The company runs on a membership model in the healthcare space, with a lean, fully remote team that ships fast and automates the boring stuff. They are looking for an operator to be the Founder's right hand: someone who gives him his time back and gives him numbers he can trust. WHY THIS ROLE IS DIFFERENT Direct line to the Founder. You work shoulder to shoulder with the CEO, not three layers down. Your judgment shapes real decisions. You build it, you don't just run it. From day one you will automate, redesign and own the systems the company runs on. A seat that grows with the company. This is the on-ramp to a broader leadership role as they scale. Impact you can see this week. Small team, no bureaucracy. What you build ships now. WHAT SUCCESS LOOKS LIKE IN YEAR ONE Month-end close done within 5 business days, every month, plus a one-page financial snapshot the CEO can act on Recurring revenue (MRR) reconciled between Stripe and HubSpot to the dollar, through a documented, repeatable process Payroll, AP and member billing running on time, every cycle A Notion project dashboard the whole team trusts, updated weekly without the CEO chasing anyone At least three recurring processes automated, with the hours saved to prove it A lighter CEO: his follow-ups, inbox triage and cross-project coordination run through you WHAT YOU WILL OWN Chief of Staff and Project Operations Own the project dashboard and pull weekly updates from the team Drive the operating cadence: priorities visible, commitments and deadlines kept Plan and run leadership QBRs: agenda, pre-work and follow-through Executive Support CEO inbox: triage, draft and manage so nothing falls through the cracks Cross-project coordination and special projects, owned end to end Be the reason things actually get done after the meeting ends Financial Operations Books and monthly close (categorization, reconciliations) with the outside CPA firm MRR reconciliation between billing and CRM, building a cleaner process AP and AR: vendor bills paid on time, member billing and collections through Stripe A simple, reliable rolling cash-flow forecast People Operations Bi-weekly payroll in Gusto Onboarding and offboarding: systems, tools and access Monthly sales commissions flowing correctly into payroll WHO YOU ARE An operator, not just an executor. You see what needs doing and do it before being asked Financially fluent: hands-on experience with month-end close, reconciliations, AP/AR and payroll (QuickBooks preferred; subscription revenue is a plus) Excellent written and spoken English. You will manage the CEO's inbox and work with the US team daily Systems-minded: you reach for automation and structure, not more manual effort Relentlessly organized. You never drop a ball and you close every loop Trustworthy and discreet. You will handle finances, payroll and the CEO's inbox BONUS POINTS Hands-on with Notion, HubSpot and Stripe, plus AI tools (Claude) and automation platforms (Zapier, Make) Prior experience as Chief of Staff, EA or operations lead supporting a founder Background in SaaS, healthcare or procurement THE DETAILS 100% remote, open to candidates based in Brazil only Long-term, high-trust role with real room to grow Finalists complete two short, paid case studies
Posted 9/22/2026
Role Overview We are seeking Spanish-speaking photographers and annotators to create evaluation materials using cropped, occluded, or degraded images of authentic real-world places, products, and media. These images should require multimodal agents to conduct multi-step web research using tools such as Google Search, Google Lens, or reverse image search before accurately generating or restoring missing content. Annotators will source real, verifiably licensed photographs, obscure key identifying details, and prepare a complete deliverable package—including prompts, grading rubrics, and solution notes—for each instance. Job responsibilities: Source authentic, high-quality photographs of real-world places, products, and media relevant to the assigned locale. Verify that each image is copyright-free, publicly available, explicitly licensed, or original, and document its license type and provenance. Conduct multi-step research using Google Search, Google Lens, reverse image search, and other reliable sources to identify and verify image subjects. Create cropped, occluded, masked, or degraded image versions that conceal key identifying details while preserving realistic visual clues. Develop clear, time-bound prompts in English without revealing the target identity or answer. Prepare concise, objective, one-sentence grading criteria for each evaluation instance. Write solution notes documenting the ground-truth identity, research process, supporting evidence, and expected outcome. Key Qualifications: Language: Native or highly fluent in Spanish, with strong working proficiency in English. All prompts, rubrics, and solution notes must be written in English. Local cultural knowledge: Strong familiarity with regionally significant but non-touristy landmarks, universities, municipal buildings, and domestic film and media within the assigned locale. This is a critical requirement. Research skills: Comfortable conducting iterative, multi-step investigative research using Google Search, Google Lens, and reverse image search to verify a ground-truth subject from partial visual clues. Image sourcing and copyright literacy: Able to source and verify copyright-free or explicitly licensed imagery from platforms such as Wikimedia Commons, Smithsonian Open Access, Creative Commons, and public-domain collections, or use original photography. Candidates must be able to document the license type and image provenance. Basic image-editing skills: Able to crop, mask, or otherwise obscure source images—or create annotated crop and mask overlays—while preserving realistic, non-hallucinated secondary visual clues. Attention to detail: Able to write concise, objective, one-sentence rubric criteria and appropriately time-bound prompts without revealing the target identity in generative or text-based prompts. Tool proficiency: Comfortable working with shared Google Drive folders, file-naming conventions, and structured deliverable submissions. Education and Experience Bachelor’s degree or equivalent practical experience in any field. Experience in AI evaluation, data annotation, content review, quality assurance, or a related analytical role is preferred but not required. Offer Details: Remote Employment Type: Contractor assignment Commitment Required: Full-time (40 hours per week), with at least 4 hours of overlap during PST business hours. Location: Argentina, Peru, Mexico, Brazil, Chile, Colombia Engagement Length: 16 weeks for Mexico, Peru, Chile, Colombia | 4 weeks for Argentina, Brazil Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/18/2026
We are seeking Portuguese-speaking photographers and annotators to create evaluation materials using cropped, occluded, or degraded images of authentic real-world places, products, and media. These images should require multimodal agents to conduct multi-step web research using tools such as Google Search, Google Lens, or reverse image search before accurately generating or restoring missing content. Annotators will source real, verifiably licensed photographs, obscure key identifying details, and prepare a complete deliverable package—including prompts, grading rubrics, and solution notes—for each instance. Job responsibilities - Source authentic, high-quality photographs of real-world places, products, and media relevant to the assigned locale. Verify that each image is copyright-free, publicly available, explicitly licensed, or original, and document its license type and provenance. Conduct multi-step research using Google Search, Google Lens, reverse image search, and other reliable sources to identify and verify image subjects. Create cropped, occluded, masked, or degraded image versions that conceal key identifying details while preserving realistic visual clues. Develop clear, time-bound prompts in English without revealing the target identity or answer. Prepare concise, objective, one-sentence grading criteria for each evaluation instance. Write solution notes documenting the ground-truth identity, research process, supporting evidence, and expected outcome. Key Qualifications Language: Native or highly fluent in Portuguese, with strong working proficiency in English. All prompts, rubrics, and solution notes must be written in English. Local cultural knowledge: Strong familiarity with regionally significant but non-touristy landmarks, universities, municipal buildings, and domestic film and media within the assigned locale. This is a critical requirement. Research skills: Comfortable conducting iterative, multi-step investigative research using Google Search, Google Lens, and reverse image search to verify a ground-truth subject from partial visual clues. Image sourcing and copyright literacy: Able to source and verify copyright-free or explicitly licensed imagery from platforms such as Wikimedia Commons, Smithsonian Open Access, Creative Commons, and public-domain collections, or use original photography. Candidates must be able to document the license type and image provenance. Basic image-editing skills: Able to crop, mask, or otherwise obscure source images—or create annotated crop and mask overlays—while preserving realistic, non-hallucinated secondary visual clues. Attention to detail: Able to write concise, objective, one-sentence rubric criteria and appropriately time-bound prompts without revealing the target identity in generative or text-based prompts. Tool proficiency: Comfortable working with shared Google Drive folders, file-naming conventions, and structured deliverable submissions. Education and Experience Bachelor’s degree or equivalent practical experience in any field. Experience in AI evaluation, data annotation, content review, quality assurance, or a related analytical role is preferred but not required. Offer Details Remote Employment Type: Contractor assignment Location: Brazil Engagement Length: 16 weeks Commitment Required: Full-time (40 hours per week), with at least 4 hours of overlap during PST business hours. Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/17/2026
About the role We are hiring 3 Senior SAP Functional Consultants for an enterprise-scale SAP S/4HANA data migration program. Each position focuses on a separate SAP functional area: Position 1 – SAP PM Functional Consultant – Data Migration Own the migration of Plant Maintenance data, including equipment, functional locations, BOMs, task lists, maintenance plans, measuring points, notifications, and work orders. Position 2 – SAP MM/EWM Functional Consultant – Data Migration Own the migration of materials, procurement, inventory, supplier, and warehouse data, including material/vendor master, purchasing documents, valuation, batches, warehouse structures, bins, handling units, and stock. Position 3 – SAP SD Functional Consultant – Data Migration Own the migration of customer and order-to-cash data, including Business Partner/customer master, sales orders, returns, deliveries, billing documents, pricing conditions, and customer-material records. Across all 3 positions, consultants will partner with migration engineers, SAP configuration teams, and business data owners throughout mapping, mock loads, validation, reconciliation, SIT/UAT, production cutover, and hypercare. Required Skills 7+ years of functional SAP experience in the relevant module: PM, MM/EWM, or SD. 2+ full S/4HANA data migration lifecycles with functional ownership of migration objects. Hands-on experience with source-to-target mapping (STTM), transformation/value-mapping rules, data quality, validation, reconciliation, and sign-off. Experience executing multiple mock loads, dress rehearsals, and production cutovers. Working knowledge of SAP S/4HANA Migration Cockpit (LTMC/LTMOM) and migration mechanisms such as BAPIs, function modules, APIs, and IDocs. Strong ability to analyze migration errors and distinguish data defects, mapping issues, configuration gaps, and tooling issues. Experience defining object dependencies, load sequencing, migration scope, and functional acceptance criteria. Strong stakeholder management and communication skills across business, SAP functional, technical, and migration teams Bonus Skills Experience with SNP CrystalBridge, Syniti ADM, SAP Data Services (BODS), Winshuttle, or comparable migration tools. Familiarity with SAP Activate methodology. Basic ABAP and SQL knowledge for troubleshooting and reconciliation. SAP/S/4HANA certification. Experience in regulated or audit-heavy enterprise environments. Exposure to AI-enabled automation or modern data migration tooling. Previous customer-facing consulting or delivery experience. Offer Details: Remote Full-time dedication (40 hours/week) Location: Brazil Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/17/2026
About the role We are an AI infrastructure company building intelligent automation solutions that help enterprises streamline complex business workflows across major enterprise platforms. Backed by leading investors and driven by an experienced founding team, we develop cutting-edge multi-agent AI systems that operate at production scale. As a Cloud Infrastructure / DevOps Engineer, you will play a key role in designing, building, and operating the infrastructure that powers AI applications across multiple cloud providers. You will own critical infrastructure components, improve platform reliability and security, and build automation that enables engineering teams to deliver software efficiently while meeting enterprise-grade security and compliance requirements.. Required Skills: 8+ years of experience as a Cloud Infrastructure or DevOps Engineer. Strong hands-on experience with Terraform. Strong hands-on experience with Kubernetes. Experience working with one or more major cloud providers, including Azure and GCP. Proven experience designing, implementing and maintaining production-grade cloud infrastructure. Experience building and managing CI/CD pipelines, release automation, GitOps workflows and pipeline-as-code practices. Experience creating automation frameworks for DevOps and developer workflows using scripting languages such as Python, Bash or similar. Strong understanding of networking fundamentals. Experience deploying, monitoring and operating highly available production systems. Ability to design reusable infrastructure and automation that works consistently across multiple customer environments. Experience collaborating closely with software engineering teams to improve delivery velocity. Strong ownership mindset with the ability to make and defend infrastructure architecture decisions. Bonus Skills: Experience with FluxCD, ArgoCD, or other GitOps controllers. Knowledge of cloud security and DevSecOps practices, including security gates, SBOM generation, and compliance automation. Experience building self-service developer tooling and operational automation. Startup and/or large technology company experience. Interest in creating clean, standardized, and portable infrastructure across cloud providers. Offer Details: Remote Full-time - 40 hours/week with overlap of 4 hours with PST Location: Brazil Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/16/2026
This is a long-term engagement with a US-based client in the employee benefits technology space. Openings on this requisition: 20 Location: Remote, anywhere in Brazil Engagement: Long-term project with a US-based client About the role As a Software Engineer, you will work in a team that partners closely with Product Managers and customers to learn about the challenges employers face while navigating the competitive employee benefits landscape. You will design solutions that solve problems in ways our customers love and work for our business. You will build the highest quality software in the latest technologies and test driven development practices. Responsibilities Collaborate with stakeholders to learn about our customers biggest challenges. Measure, inspect, and drive decisions using data. Design, test, code, and instrument new solutions. Strengthen and drive our engineering process with TDD and BDD, Microservice and Vertical Slice Architectures. Support live applications, promote proactive monitoring, rapid incident response, and continuous improvement. Analyze existing systems and processes to identify bottlenecks and opportunities for improvements. Mentor and learn from your peers, foster continuous learning within your team and organization. Become a subject matter expert in your functional area and best practices. Assess unique circumstances and apply creative problem-solving techniques to resolve issues or suggest various approaches. Independently complete work and proactively review with others. Highly self-motivated, take ownership of your work, actively seek out ways to contribute, and require minimal supervision. Required Qualifications Bachelor's degree in Computer Science, Software Engineering, or related field, OR demonstrable equivalent experience. At least 2 years of experience in software engineering (2 to 5 years). Strong problem-solving and analytical skills. Excellent communication and collaboration skills. Passionate about keeping up with modern technologies and design. Strong proficiency in Angular and/or React. Experience building and consuming REST APIs. Proven track record of writing comprehensive unit tests and test suites. Strong understanding of software security principles and OWASP guidelines. Proficiency with Git version control and CI/CD pipelines. Experience with Agile development methodologies. Track record of delivering complex projects on schedule. Experience in writing performant stored procedures and functions. Experience in developing Cloud-Native applications and services. Technology Must-Haves C# Docker Modern RDBMS (MS SQL, Postgres, MySQL) ASP.NET RESTful API design Event-driven design with Kafka or an equivalent messaging platform (Azure Service Bus, RabbitMQ, SQS) Modern Web UI frameworks and libraries (Angular, React) Technology Nice-To-Have or Dedicate to Learning Quickly Kubernetes Helm / ArgoCD Terraform GitHub Actions (the CI/CD pipeline you will work with) NoSQL databases GraphQL Azure
Posted 9/16/2026
Role Overview We are seeking experienced AI Evaluation Engineers (Engineering Simulation & Design) to author and validate "model-breaking," simulation-based engineering design problems to train and evaluate state-of-the-art AI agents. Operating across major engineering disciplines—including Electrical, Mechanical, Control Systems, Aerospace, Systems, and Robotics—you will create complex, multi-constraint tasks where AI agents must interpret requirements, navigate trade-offs, configure open-source simulation tools, diagnose failures, and iterate toward valid solutions. You will analyze agent execution logs, expose systemic reasoning gaps, and build automated, objective graders to elevate frontier model performance. Job Requirements: Education & Expertise: Master’s degree or PhD in Electrical, Mechanical, Aerospace, with 10+ years of hands-on engineering design experience. Simulation Tooling: Proficiency with at least one domain-relevant open-source simulation package (e.g., ngspice, PySpice, OpenFOAM, FEniCSx, CalculiX, python-control, CadQuery, build123d, OpenModelica, Cantera, Gmsh) combined with strong Python scripting skills. AI Evaluation & Failure Diagnostics: Hands-on experience with modern LLMs/coding agents and evaluation concepts (pass@k, failure-mode analysis, nondeterministic behavior), with the ability to audit trajectory logs and isolate core reasoning/tool-use failures. Domain Rigor & Precision: Uncompromising attention to physical plausibility, unit consistency, boundary conditions, convergence criteria, and technical documentation. Availability & Commitment: Talent must have weekend on-call availability (part-time engagement is acceptable). Technical Infrastructure: Personal desktop/laptop equipped with a stable, high-speed internet connection in a remote setup. Job Responsibilities Model-Breaking Problem Design: Author original, self-contained engineering design tasks with competing constraints, explicit optimization targets, validated reference solutions, and objective autograders. Environment & Simulation Integration: Build, run, and validate problem environments using open-source simulation tools and custom Python test benches. Trajectory Analysis & Failure Mode Taxonomy: Evaluate coding agent outputs and execution logs across repeated trials to identify systemic failure modes (e.g., misinterpreting simulator feedback, premature design convergence, physically impossible geometries). Difficulty Calibration & Benchmark Refinement: Iteratively refine problem difficulty based on empirical model performance data without introducing ambiguity or missing information. Cross-Functional Collaboration: Partner with AI researchers, pod leads, and domain experts to integrate high-rigor benchmarks into the model evaluation pipeline. Domains: Electrical Engineering Mechanical Engineering Aerospace Engineering Education & Experience Bachelor's degree or equivalent practical experience in any field. Experience in AI evaluation, data annotation, content review, quality assurance, or a related analytical role is preferred but not required. Offer Details: Remote 40 hours per week with 4 hours of overlap with PST Location: Bangladesh, India, Indonesia, Vietnam, Egypt, Ghana, Turkey, Brazil, Colombia Engagement Type: Short Term Contract (24 weeks) Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/15/2026
Role Overview Consultants will review, evaluate, and refine prompts and PowerPoint presentations to produce accurate, high-quality enterprise data for clients. The role requires applying strong business judgment, structured thinking, and presentation expertise to ensure content is clear, relevant, logically organized, and aligned with client objectives and professional standards. Requirements: Minimum 3+ years of experience creating and working with presentations in management consulting, corporate strategy, corporate finance, or similar professional environments. Strong expertise in developing, editing, and reviewing polished, enterprise-level PowerPoint presentations for business and executive audiences. Exceptional written communication skills, with the ability to present complex information clearly, concisely, and professionally. Proven ability to review, edit, rewrite, and refine slide content while maintaining accuracy, consistency, and the intended message. Responsibilities: Strong expertise in creating, editing, formatting, and reviewing enterprise-level PowerPoint presentations aligned with professional standards and brand guidelines. Exceptional attention to detail, ensuring consistency in layout, typography, colors, visual hierarchy, alignment, and overall presentation quality. Ability to transform complex information into clear, visually compelling, and well-structured slides for business and executive audiences. Proven ability to work efficiently under tight deadlines while maintaining accuracy, consistency, and a high standard of quality. Education & Experience: Bachelor's degree or equivalent practical experience in any field. Experience in AI evaluation, data annotation, content review, quality assurance, or a related analytical role is preferred but not required. Offer Details: Remote Full-time dedication: 40 hours/week with 4 hours of overlap with PST. Location: Bangladesh, India, Indonesia, Vietnam, Egypt, Ghana, Turkey, Brazil, Colombia Engagement Length: 10 weeks Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/9/2026
Role Overview We're building one of the most rigorous STEM AI training datasets in the industry. The SciCode project involves creating high-quality scientific coding tasks that are used to train and evaluate frontier AI models. As a SciCode Trainer, you will be directly contributing to cutting-edge AI research by authoring, implementing, and reviewing complex scientific problems across core STEM disciplines. Responsibilities: Write scientific problem specifications consisting of one main problem and a minimum of 3 sub-problems, all logically connected and progressively building toward the main problem solution Implement verified golden solutions in Python with complete unit test coverage Develop discriminative test cases that clearly distinguish between correct and incorrect model outputs. Perform quality control (QC) checks on the Central Task Platform (CTP), including Level 1 structural checks and Level 2 quality criteria. Refine tasks based on QC feedback to meet Pass@K evaluation criteria across multiple language models (LLMs) acting as judges (GPT, Gemini, Nemotron). Maintain high output quality with a low rework rate, targeting consistent L1 approval on first submission Participate in sync calls for reviews, feedback sessions, and project standups during overlap hours Required Qualifications: Master's or PhD in Material science or related field. Strong Python programming skills with experience in scientific computing Ability to write rigorous, well-posed scientific problems with clear constraints and expected outputs Attention to detail - tasks must meet strict rubrics for well-posedness, test case discriminativeness, scientific correctness, and determinism Prior experience in AI data annotation, research, or scientific writing Familiarity with LLM evaluation frameworks or coding benchmarks Experience with libraries such as NumPy, SciPy, SymPy, or domain-specific scientific tools. Published research or academic project experience in a STEM domain Quality Standards Offer Details: Remote Commitments Required: Full-time dedication (40 hours/week) - Overlap of 4 hours with PST Location: Bangladesh, Brazil, Colombia, Egypt, Ghana, India, Indonesia, Turkey, Vietnam Engagement Length: 8 weeks About the evaluation process: No interviews or assessments are required. The selection process consists of reviewing your CV and Google Scholar profile link to evaluate your academic and professional background. Please note: This position is paid on a task based model. The hourly rate listed in this form is only an estimated average to provide compensation context. The final payment will depend on the tasks completed. Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/8/2026
Role Overview We're building one of the most rigorous STEM AI training datasets in the industry. The SciCode project involves creating high-quality scientific coding tasks that are used to train and evaluate frontier AI models. As a SciCode Trainer, you will be directly contributing to cutting-edge AI research by authoring, implementing, and reviewing complex scientific problems across core STEM disciplines. Responsibilities Write scientific problem specifications consisting of one main problem and a minimum of 3 sub-problems, all logically connected and progressively building toward the main problem solution Implement verified golden solutions in Python with complete unit test coverage Develop discriminative test cases that clearly distinguish between correct and incorrect model outputs. Perform quality control (QC) checks on the Central Task Platform (CTP), including Level 1 structural checks and Level 2 quality criteria. Refine tasks based on QC feedback to meet Pass@K evaluation criteria across multiple language models (LLMs) acting as judges (GPT, Gemini, Nemotron). Maintain high output quality with a low rework rate, targeting consistent L1 approval on first submission Participate in sync calls for reviews, feedback sessions, and project standups during overlap hours Required Qualifications Master's or PhD in Biology. Strong Python programming skills with experience in scientific computing Ability to write rigorous, well-posed scientific problems with clear constraints and expected outputs Attention to detail - tasks must meet strict rubrics for well-posedness, test case discriminativeness, scientific correctness, and determinism Prior experience in AI data annotation, research, or scientific writing Familiarity with LLM evaluation frameworks or coding benchmarks Experience with libraries such as NumPy, SciPy, SymPy, or domain-specific scientific tools. Published research or academic project experience in a STEM domain Quality Standards Offer Details: Remote Commitments Required: Full-time dedication (40 hours/week) - Overlap of 4 hours with PST Location: Bangladesh, Brazil, Colombia, Egypt, Ghana, India, Indonesia, Turkey, Vietnam Engagement Length: 8 weeks About the evaluation process: No interviews or assessments are required. The selection process consists of reviewing your CV and Google Scholar profile link to evaluate your academic and professional background. Please note: This position is paid on a task based model. The hourly rate listed in this form is only an estimated average to provide compensation context. The final payment will depend on the tasks completed. Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/8/2026
Role Overview We're building one of the most rigorous STEM AI training datasets in the industry. The SciCode project involves creating high-quality scientific coding tasks that are used to train and evaluate frontier AI models. As a SciCode Trainer, you will be directly contributing to cutting-edge AI research by authoring, implementing, and reviewing complex scientific problems across core STEM disciplines. Responsibilities: Write scientific problem specifications consisting of one main problem and a minimum of 3 sub-problems, all logically connected and progressively building toward the main problem solution Implement verified golden solutions in Python with complete unit test coverage Develop discriminative test cases that clearly distinguish between correct and incorrect model outputs. Perform quality control (QC) checks on the Central Task Platform (CTP), including Level 1 structural checks and Level 2 quality criteria. Refine tasks based on QC feedback to meet Pass@K evaluation criteria across multiple language models (LLMs) acting as judges (GPT, Gemini, Nemotron). Maintain high output quality with a low rework rate, targeting consistent L1 approval on first submission Participate in sync calls for reviews, feedback sessions, and project standups during overlap hours Required Qualifications: Master's or PhD in Chemistry. Strong Python programming skills with experience in scientific computing Ability to write rigorous, well-posed scientific problems with clear constraints and expected outputs Attention to detail - tasks must meet strict rubrics for well-posedness, test case discriminativeness, scientific correctness, and determinism Prior experience in AI data annotation, research, or scientific writing Familiarity with LLM evaluation frameworks or coding benchmarks Experience with libraries such as NumPy, SciPy, SymPy, or domain-specific scientific tools. Published research or academic project experience in a STEM domain Quality Standards Offer Details: Remote Commitments Required: Full-time dedication (40 hours/week) - Overlap of 4 hours with PST Location: Bangladesh, Brazil, Colombia, Egypt, Ghana, India, Indonesia, Turkey, Vietnam Engagement Length: 8 weeks About the evaluation process: No interviews or assessments are required. The selection process consists of reviewing your CV and Google Scholar profile link to evaluate your academic and professional background. Please note: This position is paid on a task based model. The hourly rate listed in this form is only an estimated average to provide compensation context. The final payment will depend on the tasks completed. Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/8/2026
Role Overview We're building one of the most rigorous STEM AI training datasets in the industry. The SciCode project involves creating high-quality scientific coding tasks that are used to train and evaluate frontier AI models. As a SciCode Trainer, you will be directly contributing to cutting-edge AI research by authoring, implementing, and reviewing complex scientific problems across core STEM disciplines. Responsibilities Write scientific problem specifications consisting of one main problem and a minimum of 3 sub-problems, all logically connected and progressively building toward the main problem solution Implement verified golden solutions in Python with complete unit test coverage Develop discriminative test cases that clearly distinguish between correct and incorrect model outputs. Perform quality control (QC) checks on the Central Task Platform (CTP), including Level 1 structural checks and Level 2 quality criteria. Refine tasks based on QC feedback to meet Pass@K evaluation criteria across multiple language models (LLMs) acting as judges (GPT, Gemini, Nemotron). Maintain high output quality with a low rework rate, targeting consistent L1 approval on first submission Participate in sync calls for reviews, feedback sessions, and project standups during overlap hours Required Qualifications Master's or PhD in Mathematics. Strong Python programming skills with experience in scientific computing Ability to write rigorous, well-posed scientific problems with clear constraints and expected outputs Attention to detail - tasks must meet strict rubrics for well-posedness, test case discriminativeness, scientific correctness, and determinism Prior experience in AI data annotation, research, or scientific writing Familiarity with LLM evaluation frameworks or coding benchmarks Experience with libraries such as NumPy, SciPy, SymPy, or domain-specific scientific tools. Published research or academic project experience in a STEM domain Quality Standards Offer Details: Remote Commitments Required: Full-time dedication (40 hours/week) - Overlap of 4 hours with PST Location: Bangladesh, Brazil, Colombia, Egypt, Ghana, India, Indonesia, Turkey, Vietnam Engagement Length: 8 weeks About the evaluation process: No interviews or assessments are required. The selection process consists of reviewing your CV and Google Scholar profile link to evaluate your academic and professional background. Please note: This position is paid on a task based model. The hourly rate listed in this form is only an estimated average to provide compensation context. The final payment will depend on the tasks completed. Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/8/2026
Role Overview We're building one of the most rigorous STEM AI training datasets in the industry. The SciCode project involves creating high-quality scientific coding tasks that are used to train and evaluate frontier AI models. As a SciCode Trainer, you will be directly contributing to cutting-edge AI research by authoring, implementing, and reviewing complex scientific problems across core STEM disciplines. Responsibilities Write scientific problem specifications consisting of one main problem and a minimum of 3 sub-problems, all logically connected and progressively building toward the main problem solution Implement verified golden solutions in Python with complete unit test coverage Develop discriminative test cases that clearly distinguish between correct and incorrect model outputs. Perform quality control (QC) checks on the Central Task Platform (CTP), including Level 1 structural checks and Level 2 quality criteria. Refine tasks based on QC feedback to meet Pass@K evaluation criteria across multiple language models (LLMs) acting as judges (GPT, Gemini, Nemotron). Maintain high output quality with a low rework rate, targeting consistent L1 approval on first submission Participate in sync calls for reviews, feedback sessions, and project standups during overlap hours Required Qualifications Master's or PhD in Physics. Strong Python programming skills with experience in scientific computing Ability to write rigorous, well-posed scientific problems with clear constraints and expected outputs Attention to detail - tasks must meet strict rubrics for well-posedness, test case discriminativeness, scientific correctness, and determinism Prior experience in AI data annotation, research, or scientific writing Familiarity with LLM evaluation frameworks or coding benchmarks Experience with libraries such as NumPy, SciPy, SymPy, or domain-specific scientific tools . Published research or academic project experience in a STEM domain Quality Standards Offer Details: Remote Commitments Required: Full-time dedication (40 hours/week) - Overlap of 4 hours with PST Location: Bangladesh, Brazil, Colombia, Egypt, Ghana, India, Indonesia, Turkey, Vietnam Engagement Length: 8 weeks About the evaluation process: No interviews or assessments are required. The selection process consists of reviewing your CV and Google Scholar profile link to evaluate your academic and professional background. Please note: This position is paid on a task based model. The hourly rate listed in this form is only an estimated average to provide compensation context. The final payment will depend on the tasks completed. Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/8/2026
We are looking for a Scientific Application Developer to support a Process Science Modelling and Data team within a global healthcare company. In this role, you will develop, enhance, test, document, and deploy scientific applications and APIs for defined project needs. You will translate agreed requirements into robust, scalable solutions, provide technical recommendations, and communicate progress, risks, and dependencies. The work will support manufacturing sites, platforms, and modalities globally, with a focus on improving process efficiency, product quality, and the usability of scientific data and models. Projects may include process capability and quality improvements, cost-of-goods initiatives, and scientific modelling needs across internal and external manufacturing sites. Key Deliverables and Responsibilities Design, develop, test, and document scientific applications and APIs that address defined manufacturing and technical requirements. Configure and use approved DevOps tools and deployment pipelines to deliver maintainable applications for manufacturing and technical users. Develop frontend web applications and interactive dashboards using a Python-based stack, including Quart (async Flask), Jinja2 templates, and Bokeh for visualization and plotting. Integrate applications with data available through the Databricks environment, including access through Databricks SQL Warehouse. Translate business and scientific requirements into intuitive, user-friendly dashboard and website experiences, with strong attention to UI/UX and data visualization. Work within established internal CI/CD tooling and deployment processes; apply sound CI/CD concepts without being responsible for standing up the pipeline from scratch. Collaborate with designated internal subject matter experts to refine requirements and integrate digital tools with approved systems and workflows. Prepare technical documentation, test evidence, deployment materials, and knowledge-transfer artifacts sufficient for internal support and maintenance. Track work against agreed milestones and promptly communicate delivery risks, technical issues, and dependencies to the designated project lead. Perform services in accordance with applicable security, data integrity, quality, regulatory, and documentation requirements. Required Experience and Capabilities Bachelor’s or master’s degree in Engineering, Computer Science, Data Science, or a related scientific or technical field, or equivalent relevant experience. Demonstrated experience developing production-quality scientific or data-centric software in Python; JavaScript experience is beneficial. Experience in UX/design for analytical or scientific applications, particularly within a Python-centric application stack Strong hands-on experience developing websites and dashboards in Python, with Quart (or Flask/async Flask) and Jinja2 or similar server-side templating approaches. Experience building interactive data visualizations and dashboards using Bokeh; experience with comparable Python visualization libraries is also relevant. Working knowledge of HTML, CSS, frontend development, and UI/UX principles, with the ability to create clear and usable interfaces for business and scientific users. Experience working with SQL and data-driven applications; familiarity with Databricks and Databricks SQL Warehouse is strongly preferred. Familiarity with CI/CD concepts and experience working within established deployment pipelines and internal DevOps tooling; candidates are not expected to build the CI/CD pipeline from scratch. Experience with software design, version control, testing, code review, documentation, and deployment practices. Working knowledge of relevant DevOps technologies such as Docker, Kubernetes, Git, and Jenkins or comparable CI/CD platforms. Ability to work independently within an agreed scope, manage priorities, estimate effort, and deliver high-quality work with limited day-to-day direction. Strong analytical, troubleshooting, and stakeholder communication skills. Experience in pharmaceutical, biotechnology, chemical, or regulated manufacturing environments is preferred; familiarity with GxP expectations is beneficial. Offer Details Remote Full-time dedication (40 hours/week) Initial 3-month contract, with the possibility of renewal. Location: Brazil, Mexico Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/8/2026
Role Overview We are looking for a Full Stack Developer with strong frontend expertise to build modern, scalable web applications. The role requires hands-on experience with Next.js on the frontend and FastAPI on the backend, along with the ability to translate Figma designs into high-quality, responsive user interfaces. The ideal candidate is detail-oriented, technically strong, and comfortable owning features end to end. Key Responsibilities Develop and maintainfrontend applications using Next.js, ensuring performance and usability. Build and integratebackend APIs using FastAPI. Convert Figma designs into pixel-perfect, responsive UI components. Collaborate closely with designers, product managers, and engineers. Ensure best practices infrontend architecture, performance, and accessibility. Participate in code reviews, debugging, and testing. Required Skills & Qualifications 4–6 years of experienceas a Full Stack Developer with frontend focus. Strong hands-on experience with Next.js (React-only experience isnot preferred). Backend development experience using FastAPI (Python). Proven experience working with Figma for UI design-to-development. Strong knowledge ofHTML, CSS, JavaScript/TypeScript. Experience working withRESTful APIs. Offer Details Remote Full-time dedication (40 hours/week) Location: Brazil, Mexico Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/8/2026
About the Role Build the technical implementation of a system that captures GitHub Copilot CLI agent traces across developer machines, scrubs sensitive content centrally in Azure, and serves a sanitized, queryable dataset for review. You’ll own the pipeline end-to-end, from the endpoint capture script to the Azure ingestion, scrub, and viewer. You are a fit: If you can take a written design and ship the working pipeline mostly on your own, and you treat privacy as a hard requirement, not an afterthought. What does day-to-day look like: Stand up the OpenTelemetry capture path: endpoint env-var/.bashrc setup, OTLP export, and a collector pipeline (receivers → processors → exporters). Build the central ingestion + privacy scrub block in Azure (secret + PII detection/redaction) before any data reaches a reviewer. Wire up the sanitized stores and a read-only trace viewer, with task↔session↔trace correlation. Required Skills: Bachelor’s degree in Computer Science, Engineering, or a related field — or equivalent practical experience. 5+ years in software / platform / data engineering, shipping production systems. 3+ years hands-on with Azure (or a comparable major cloud with clear transfer to Azure PaaS), including 2+ years building observability or telemetry pipelines with direct OpenTelemetry experience — OTLP, collectors, spans/traces, semantic conventions. Azure PaaS: Container Apps/AKS, API Management, ADX/Kusto or App Insights, ADLS Gen2/Blob, Key Vault, Entra ID. Solid Python or Node.js, plus Bash scripting. Practical data privacy / PII & secret scrubbing experience (regex, entropy, redaction/tokenization). Data-pipeline and cloud-security fundamentals (RBAC, private networking, audit). Nice to Have: Grafana / telemetry-backend setup; LLM-agent or GenAI observability exposure; cross-cloud identity federation (AAD ⇄ GCP IAM). Offer Details: Remote Location: USA and LATAM Engagement Type: Short Term Contract (3 months) At least 4 hours per day and minimum 20 hours per week with overlap of 4 hours with PST. Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/7/2026
Empresa: confidencial (startup recém-capitalizada) Área: Engenharia, founding team Local: SP Capital, 100% presencial Contratação: PJ + equity Sobre a empresa Nosso cliente é uma startup que acabou de captar investimento. O time ainda é pequeno, os fundadores estão na operação e o produto está sendo construído agora. Quem entra nesse momento não entra como funcionário número 30: entra como parte do founding team, com equity e com voz nas decisões que vão definir o produto e a empresa. A vaga A pessoa vai ser a referência técnica do time nessa fase inicial. Na prática, isso significa: Construir o produto de ponta a ponta: backend, frontend, infra e integrações Tomar decisões de arquitetura pensando em velocidade agora e escala depois Trabalhar lado a lado com os fundadores, todo dia, no mesmo escritório Lidar com escopo que muda conforme o mercado responde Ajudar a montar o time técnico que vem depois O que precisam 8+ anos de experiência em engenharia de software, atuando hoje como Senior, Staff ou Principal Engineer Perfil técnico generalista: você já transitou entre backend, frontend e infra e resolve o que aparece na frente Experiência prévia em fintech é diferencial forte Perfil empreendedor: já fundou algo, já foi um dos primeiros de uma startup ou tocou projeto próprio com dono do resultado Conforto com ambiguidade e com fase de zero a um Disponibilidade para trabalhar 100% presencial em SP O que oferecem Contratação PJ com equity: participação real em uma empresa recém-capitalizada Posição de founding team, com peso nas decisões técnicas e de produto Time pequeno e contato direto com os fundadores OBS: Na próxima página vai perguntar nível de inglês, mas é padrão do nosso ATS. Não é necessário inglês para essa vaga.
Posted 9/5/2026
Role Overview We are seeking a FreeCAD Task Executor with practical experience in mechanical design and CAD workflows. The role involves executing complex modeling tasks, recording human reference demonstrations, validating task solvability, identifying design edge cases, and testing automated CAD verifiers. Target Applications and Toolchain FreeCAD, Part Design, Sketcher, TechDraw, Assembly, and STEP/IGES export. Job Responsibilities Execute multi-step parametric CAD tasks in FreeCAD using remote Ubuntu/Linux desktop environments. Record clean human execution trajectories, including mouse clicks, sketch constraints, dimensions, and input parameters. Identify ambiguous design briefs, under-constrained sketches, missing library models, and FreeCAD stability issues. Execute alternative valid modeling paths to verify that automated grading scripts accept technically correct designs. Job Requirements Practical hands-on experience with FreeCAD for 3D part modeling and technical drawing. Strong knowledge of parametric CAD workflows, sketches, constraints, dimensions, and assemblies. Ability to navigate CAD software efficiently and troubleshoot sketch constraints. High attention to dimensional accuracy and design requirements. Clear written English for submitting execution logs and solvability feedback. Experience working in remote Linux desktop environments, including noVNC or VNC is preferred Experience in CAD quality assurance, technical testing, or design benchmarking is preferred. Engagement Details: Engagement Type: Pay per task AHT (Average Handling Time) : 60 minutes per task. Start Date: As soon as possible. Duration: 5 weeks, with potential for extension. Commitment Required: Full-time. Location: Bangladesh, Brazil, Colombia, Egypt, Ghana, India, Indonesia, Kenya, Turkey, Vietnam Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/4/2026
Role Overview We are seeking a KiCad Task Executor with practical experience in electronic design automation, schematic capture, and multi-layer PCB layout. The role involves executing complex hardware-design tasks, recording human reference demonstrations, validating task solvability, identifying electrical edge cases, and testing automated grading rules. Target Applications and Toolchain KiCad, Eeschema, Pcbnew, Gerber, Design Rule Check (DRC), and the KiCad Python API. Job Responsibilities Complete multi-step schematic-wiring and PCB trace-routing tasks in remote Ubuntu/Linux desktop environments. Record clean human execution trajectories, including keystrokes, mouse-routing actions, footprint placement, and DRC runs. Identify ambiguous electrical specifications, missing footprints, clearance conflicts, and software-interface issues. Route PCBs and wire schematics using alternative valid design paths to verify accurate automated grading. Job Requirements Practical hands-on experience with KiCad for schematic capture and multi-layer PCB routing. Strong knowledge of PCB layouts, schematic wiring, footprints, symbols, routing, and DRC error resolution. Ability to navigate EDA tools efficiently and apply routing best practices. High attention to pin assignments, component orientation, and clearance constraints. Clear written English for submitting bug reports and execution feedback. Experience working in remote Linux desktop environments, including noVNC or VNC is preferred Experience in PCB layout review, hardware testing, or EDA benchmarking is preferred. Engagement Details: Engagement Type: Pay per task AHT (Average Handling Time) : 60 minutes per task. Start Date: As soon as possible. Duration: 5 weeks, with potential for extension. Commitment Required: Full-time. Location: Bangladesh, Brazil, Colombia, Egypt, Ghana, India, Indonesia, Kenya, Turkey, Vietnam Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/4/2026
Role Overview Seeking a Senior Backend Engineer to design and build the automated GCP billing pipeline for the UC-003 Billing Prep use case. This role transitions a manual, Excel-based process into a rules-driven, AI-augmented platform integrated with GCP infrastructure and the Impact billing system. Required Qualifications (Must Haves): 5+ years backend engineering experience with Python or Go in production settings. Deep technical capability in GCP core services: Cloud Run, Cloud SQL (PostgreSQL), Cloud Scheduler, Pub/Sub, Cloud Build, Artifact Registry, Cloud Monitoring, and Cloud Armor. Hands-on expertise building rules engines, configurable business logic frameworks, RESTful microservices, and event-driven architectures. Advanced PostgreSQL proficiency (schema design, indexing, transaction optimization) and CI/CD / DevOps practices. Prior background integrating with ERP/financial billing environments (PeopleSoft, SAP, Oracle) and processing financial data pipelines (charge reconciliation, accounts payable). Good to Have: GCP Professional Cloud Developer or Professional Data Engineer certification. Real estate financial operations domain experience. Exposure to workflow orchestration tools (Cloud Workflows, Prefect) and calling LLM/AI services (Vertex AI SDK) from backend systems. Key Responsibilities: Rules Engine & Pipeline Development: Build GCP-native rules engines for BIL file processing (PeopleSoft data classification, PO reconciliation, charge code splits) and expose REST APIs on Cloud Run. Data Ingestion & Integration: Create automated ingestion pipelines using Cloud Scheduler and Pub/Sub, and build resilient, bidirectional integrations with the Impact billing system. Database Architecture & State: Design PostgreSQL schemas on Cloud SQL, implementing full audit logging, data versioning, and batch rollback capabilities. Platform Operations & CI/CD: Establish Cloud Build automated testing gates, implement structured monitoring/alerting, enforce Cloud Armor security controls, and support UAT through go-live. Offer Details: Location: India - Hyderabad (On-site) Engagement Type: Full-time Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/3/2026
Role Overview Work on projects focused on fine-tuning and improving large language models such as ChatGPT. Analyzed complex content, summarized key information, validated claims through online research, and solved logical problems. Created high-quality analytical responses, detailed explanations, and training scenarios to enhance model accuracy. Developed expertise in leveraging AI to strengthen research and analytical capabilities. Key Responsibilities: Analyze complex content, datasets, logical puzzles, and research-based questions. Research online sources to verify claims and validate information accuracy. Summarize large volumes of content into concise and logical sections. Evaluate AI-generated responses and provided accurate corrections with explanations. Create analytical scenarios and annotations to improve large language model performance. Requirements: Recent graduates or currently pursuing a degree Demonstrate fluency in the U.S. English (en-US) Maintain an active, previously established Gemini App account with genuine conversation history (Avoid creating an account solely for participation in this project) Maintain at least 10 separate conversations related to education, learning, academics, exam preparation, research, projects, or skill development (This will be validated before joining). Apply strong research, analytical, logical reasoning, and problem-solving abilities. Provide constructive feedback supported by detailed explanations and annotations. Apply creative and lateral thinking to solve unfamiliar and complex problems. Work independently while maintaining productivity in a remote environment Preferred Skills: Pursued or completed a bachelor’s degree in Engineering, Literature, Journalism, Communications, Arts, Statistics, or a related field. Demonstrated relevant professional experience in analysis, research, writing, editing, translation, or journalism. Used Microsoft Excel and Google Workspace for documentation and data-related tasks. Interpreted data and applied basic arithmetic to evaluate trends and relationships. Applied logical reasoning and AI-assisted tools to improve research and analytical outcomes. Offer Details: Remote Full-time dedication Location: United States Engagement Length: 16 weeks Employment type: Contractor assignment Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/2/2026
We are looking for an experienced Senior Salesforce Administrator to join a fast-paced Go-to-Market Systems team supporting Sales and Operations. In this role, you will serve as the first point of contact for Salesforce users, ensuring the platform remains reliable, accessible, and efficient for day-to-day business operations. You will be responsible for resolving user issues, managing access, performing light system configuration, maintaining data quality, and collaborating with Product and Engineering teams to escalate and troubleshoot more complex platform challenges. Required Skills 5–7 years of hands-on experience as a Salesforce Administrator. Strong knowledge of Salesforce administration, including user provisioning, de-provisioning, and access management. Experience managing roles, profiles, permission sets, sharing rules, and Salesforce security models. Solid understanding of Salesforce standard objects, including Leads, Contacts, Accounts, and Opportunities. Experience troubleshooting user access, visibility issues, and standard Salesforce errors. Ability to perform routine Salesforce configuration, including fields, page layouts, list views, reports, and other declarative features. Experience supporting data quality initiatives through imports, exports, and data maintenance. Familiarity with monitoring Salesforce Flows, Validation Rules, and other standard automations. Strong analytical and problem-solving skills with the ability to identify root causes before escalating issues. Excellent written and verbal communication skills with a customer-focused mindset. Experience documenting processes, troubleshooting guides, and standard operating procedures. Bonus Skills Salesforce Certified Administrator (ADM-201). Experience supporting Go-to-Market, Sales, or Revenue Operations teams. Familiarity with ticketing and support queue management. Experience collaborating with Product and Engineering teams in cross-functional environments. Knowledge of Salesforce best practices for system administration and governance. Offer Details Remote Full-time dedication (40 hours/week) Location: Brazil, Mexico Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/2/2026
About the Role: We are looking for experienced software engineers (tech lead level) who are familiar with high-quality public GitHub repositories and can contribute to this project. You should have experience working with well-maintained, widely-used repos with 5000+ stars. This role involves hands-on software engineering work, including development environment automation, issue triaging, and evaluating test coverage and quality. What does day-to-day look like: Analyze and triage GitHub issues across trending open-source libraries. Set up and configure code repositories, including Dockerization and environment setup. Evaluating unit test coverage and quality. Modify and run codebases locally to assess LLM performance in bug-fixing scenarios. Collaborate with researchers to design and identify repositories and issues that are challenging for LLMs. Opportunities to lead a team of junior engineers to collaborate on projects. Required Skills: Strong 3+ yrs of experience with C# Proficiency with Git, Docker, and basic software pipeline setup. Ability to understand and navigate complex codebases. Comfortable running, modifying, and testing real-world projects locally. Experience contributing to or evaluating open-source projects is a plus. Nice to Have: Previous participation in LLM research or evaluation projects. Experience building or testing developer tools or automation agents. Offer details: Remote Full-time dedication (40 hours/week) - 4 hrs/day overlap with PST Location: Bangladesh, Brazil, Colombia, Egypt, Ghana, India, Indonesia, Turkey, Vietnam Engagement Length: 3 months Employment type: Contractor assignment Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/1/2026
About the Role: We are looking for experienced software engineers (tech lead level) who are familiar with high-quality public GitHub repositories and can contribute to this project. You should have experience working with well-maintained, widely-used repos with 5000+ stars. This role involves hands-on software engineering work, including development environment automation, issue triaging, and evaluating test coverage and quality What does day-to-day look like: Analyze and triage GitHub issues across trending open-source libraries. Set up and configure code repositories, including Dockerization and environment setup. Evaluating unit test coverage and quality. Modify and run codebases locally to assess LLM performance in bug-fixing scenarios. Collaborate with researchers to design and identify repositories and issues that are challenging for LLMs. Opportunities to lead a team of junior engineers to collaborate on projects. Required Skills: Strong 3+ yrs of experience with Go Proficiency with Git, Docker, and basic software pipeline setup. Ability to understand and navigate complex codebases. Comfortable running, modifying, and testing real-world projects locally. Experience contributing to or evaluating open-source projects is a plus. Nice to Have: Previous participation in LLM research or evaluation projects. Experience building or testing developer tools or automation agents. Offer details: Remote Full-time dedication (40 hours/week) - 4 hrs/day overlap with PST Location: Bangladesh, Brazil, Colombia, Egypt, Ghana, India, Indonesia, Turkey, Vietnam Engagement Length: 3 months Employment type: Contractor assignment Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/1/2026
About the Role: We are looking for experienced software engineers (tech lead level) who are familiar with high-quality public GitHub repositories and can contribute to this project. You should have experience working with well-maintained, widely-used repos with 5000+ stars. This role involves hands-on software engineering work, including development environment automation, issue triaging, and evaluating test coverage and quality. What does day-to-day look like: Analyze and triage GitHub issues across trending open-source libraries. Set up and configure code repositories, including Dockerization and environment setup. Evaluating unit test coverage and quality. Modify and run codebases locally to assess LLM performance in bug-fixing scenarios. Collaborate with researchers to design and identify repositories and issues that are challenging for LLMs. Opportunities to lead a team of junior engineers to collaborate on projects. Required Skills: Strong 3+ yrs of experience with Ruby Proficiency with Git, Docker, and basic software pipeline setup. Ability to understand and navigate complex codebases. Comfortable running, modifying, and testing real-world projects locally. Experience contributing to or evaluating open-source projects is a plus. Nice to Have: Previous participation in LLM research or evaluation projects. Experience building or testing developer tools or automation agents. Offer details: Remote Full-time dedication (40 hours/week) - 4 hrs/day overlap with PST Location: Bangladesh, Brazil, Colombia, Egypt, Ghana, India, Indonesia, Turkey, Vietnam Engagement Length: 3 months Employment type: Contractor assignment Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/1/2026
About the Role: We are looking for experienced software engineers (tech lead level) who are familiar with high-quality public GitHub repositories and can contribute to this project. You should have experience working with well-maintained, widely-used repos with 5000+ stars. This role involves hands-on software engineering work, including development environment automation, issue triaging, and evaluating test coverage and quality What does day-to-day look like: Analyze and triage GitHub issues across trending open-source libraries. Set up and configure code repositories, including Dockerization and environment setup. Evaluating unit test coverage and quality. Modify and run codebases locally to assess LLM performance in bug-fixing scenarios. Collaborate with researchers to design and identify repositories and issues that are challenging for LLMs. Opportunities to lead a team of junior engineers to collaborate on projects. Required Skills: Strong 3+ yrs of experience with Rust Proficiency with Git, Docker, and basic software pipeline setup. Ability to understand and navigate complex codebases. Comfortable running, modifying, and testing real-world projects locally. Experience contributing to or evaluating open-source projects is a plus. Nice to Have: Previous participation in LLM research or evaluation projects. Experience building or testing developer tools or automation agents. Offer details: Remote Full-time dedication (40 hours/week) - 4 hrs/day overlap with PST Location: Bangladesh, Brazil, Colombia, Egypt, Ghana, India, Indonesia, Turkey, Vietnam Engagement Length: 3 months Employment type: Contractor assignment Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/1/2026
About the Role: We are looking for experienced software engineers (tech lead level) who are familiar with high-quality public GitHub repositories and can contribute to this project. This role involves hands-on software engineering work, including development environment automation, issue triaging, and evaluating test coverage and quality. What does day-to-day look like: Analyze and triage GitHub issues across trending open-source libraries. Set up and configure code repositories, including Dockerization and environment setup. Evaluating unit test coverage and quality. Modify and run codebases locally to assess LLM performance in bug-fixing scenarios. Collaborate with researchers to design and identify repositories and issues that are challenging for LLMs. Opportunities to lead a team of junior engineers to collaborate on projects. Required Skills: Minimum 3+ years of overall experience Strong experience with at least one of the following languages: C++ Proficiency with Git, Docker, and basic software pipeline setup. Ability to understand and navigate complex codebases. Comfortable running, modifying, and testing real-world projects locally. Experience contributing to or evaluating open-source projects is a plus. Nice to Have: Previous participation in LLM research or evaluation projects. Experience building or testing developer tools or automation agents. Offer Details: Remote Full-time dedication (40 hours/week) - 4 hrs/day overlap with PST Location: Bangladesh, Brazil, Colombia, Egypt, Ghana, India, Indonesia, Turkey, Vietnam Engagement Length: 3 months Employment type: Contractor assignment Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/1/2026
About the Role We are looking for experienced software engineers (tech lead level) who are familiar with high-quality public GitHub repositories and can contribute to this project. You should have experience working with well-maintained, widely-used repos with 5000+ stars. This role involves hands-on software engineering work, including development environment automation, issue triaging, and evaluating test coverage and quality What does day-to-day look like Analyze and triage GitHub issues across trending open-source libraries. Set up and configure code repositories, including Dockerization and environment setup. Evaluating unit test coverage and quality. Modify and run codebases locally to assess LLM performance in bug-fixing scenarios. Collaborate with researchers to design and identify repositories and issues that are challenging for LLMs. Opportunities to lead a team of junior engineers to collaborate on projects. Required Skills: Strong experience with Java, Proficiency with Git, Docker, and basic software pipeline setup. Ability to understand and navigate complex codebases. Comfortable running, modifying, and testing real-world projects locally. Experience contributing to or evaluating open-source projects is a plus. Nice to Have: Previous participation in LLM research or evaluation projects. Experience building or testing developer tools or automation agents. Offer details: Full-time dedication (40 hours/week) - 4 hrs/day overlap with PST Location: Bangladesh, Brazil, Colombia, Egypt, Ghana, India, Indonesia, Turkey, Vietnam Engagement Length: 3 months Employment type: Contractor assignment Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/1/2026
About the Role We are looking for experienced software engineers (tech lead level) who are familiar with high-quality public GitHub repositories and can contribute to this project. You should have experience working with well-maintained, widely-used repos with 5000+ stars. This role involves hands-on software engineering work, including development environment automation, issue triaging, and evaluating test coverage and quality What does day-to-day look like: Analyze and triage GitHub issues across trending open-source libraries. Set up and configure code repositories, including Dockerization and environment setup. Evaluating unit test coverage and quality. Modify and run codebases locally to assess LLM performance in bug-fixing scenarios. Collaborate with researchers to design and identify repositories and issues that are challenging for LLMs. Opportunities to lead a team of junior engineers to collaborate on projects. Required Skills: Strong experience with JavaScript Proficiency with Git, Docker, and basic software pipeline setup. Ability to understand and navigate complex codebases. Comfortable running, modifying, and testing real-world projects locally. Experience contributing to or evaluating open-source projects is a plus. Nice to Have: Previous participation in LLM research or evaluation projects. Experience building or testing developer tools or automation agents. Offer details: Remote Full-time dedication (40 hours/week) - 4 hrs/day overlap with PST Location: Bangladesh, Brazil, Colombia, Egypt, Ghana, India, Indonesia, Turkey, Vietnam Engagement Length: 3 months Employment type: Contractor assignment Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/1/2026
About the Role We are looking for a Front-end Engineer to join a project with a global healthcare company. The ideal candidate will have strong front-end development experience, some exposure to back-end development, and experience taking projects from PoC to production. Required Skills 5+ years of experience Strong proficiency in front-end development, with experience in: Python, Flask, React, HTML, CSS Some back-end development experience (our backend is primarily built with LangChain) Proven ability to transition projects from PoC to production Strong problem-solving skills and attention to detail Eager to learn attitude Ability to work independently Offer Details Remote Full-time dedication (40 hours/week) Location: Brazil, Mexico Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/1/2026
About the Role: We are looking for experienced software engineers (tech lead level) who are familiar with high-quality public GitHub repositories and can contribute to this project. You should have experience working with well-maintained, widely-used repos with 5000+ stars. This role involves hands-on software engineering work, including development environment automation, issue triaging, and evaluating test coverage and quality. What does day-to-day look like: Analyze and triage GitHub issues across trending open-source libraries. Set up and configure code repositories, including Dockerization and environment setup. Evaluating unit test coverage and quality. Modify and run codebases locally to assess LLM performance in bug-fixing scenarios. Collaborate with researchers to design and identify repositories and issues that are challenging for LLMs. Opportunities to lead a team of junior engineers to collaborate on projects. Required Skills: Strong experience with Python, Proficiency with Git, Docker, and basic software pipeline setup. Ability to understand and navigate complex codebases. Comfortable running, modifying, and testing real-world projects locally. Experience contributing to or evaluating open-source projects is a plus. Nice to Have: Previous participation in LLM research or evaluation projects. Experience building or testing developer tools or automation agents. Offer details: Full-time dedication (40 hours/week) - 4 hrs/day overlap with PST Location: Bangladesh, Brazil, Colombia, Egypt, Ghana, India, Indonesia, Turkey, Vietnam Engagement Length: 3 months Employment type: Contractor assignment Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/1/2026
Role Overview We are seeking senior Infrastructure or Backend Engineers to operate and support critical GCP services, including a Perplexity-customized QC service and Cloud Run-based runner and packaging infrastructure. The role requires independent ownership of deployments, production issues, uptime, and trainer support for three months. What does day-to-day look like: Manage and monitor production deployments across GCP-based services. Maintain reliability, uptime, and performance across distributed workloads. Debug and resolve backend, infrastructure, and deployment issues. Support the runner and final-packaging systems on GCP Cloud Run. Maintain and improve the QC service customized for Perplexity. Troubleshoot failures across hundreds of concurrent runs and large pod fleets. Respond independently to trainers’ technical questions and operational requests. Collaborate with engineering and operations teams on incident resolution and service improvements. Document recurring issues, fixes, deployment procedures, and operational runbooks. Requirements 5+ years of hands-on experience in infrastructure engineering, backend engineering, or production systems. Strong production experience with Google Cloud Platform, particularly Cloud Run and related services. Advanced Python backend development experience, including debugging and maintaining production services. Strong knowledge of Kubernetes, containerized workloads, SQL databases, and distributed systems. Experience operating highly concurrent production workloads and independently handling deployments, incidents, and bug fixes. Offer Details Remote Commitments Required: 8 hours per day with a 6-8 hour overlap with PST. Location: Candidates based in the United States or Latin American time zones are preferred. Employment Type: Contractor position Duration of Contract: 3 months [expected start date is immediate]. Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 9/1/2026
Role Overview We are seeking experienced Aerospace / Flight-Dynamics Engineer to author aero/flight tasks in OpenVSP, XFLR5, FlightGear+JSBSim, OpenRocket and QGroundControl; build golden geometries/mission plans and score geometry/route/airframe files. Key Responsibilities Author aero/flight tasks and validate designs using industry-standard tools (OpenVSP, XFLR5, FlightGear+JSBSim, OpenRocket and QGroundControl) Develop realistic engineering problem statements, simulation models, calculations, schematics, and supporting technical assets. Build golden geometries/mission plans and score geometry/route/airframe files. Define inputs, assumptions, constraints, operating conditions, output metrics, and objective acceptance criteria. Validate results through analytical checks and review other engineers’ submissions for technical correctness and completeness. Translate technical requirements into simulation models and workflows. Optimize designs based on simulation results. Collaborate with cross-functional engineering teams. Document methodologies, processes, and outcomes. Key Requirements Candidates must have a minimum of 5+ years of experience working in Aerospace / Flight-Dynamics Hands-on experience with scripting Proficient in 2+ of the listed tools: OpenVSP, XFLR5, FlightGear+JSBSim, OpenRocket, QGroundControl Analytical mindset with strong problem-solving skills Ability to work independently in a fast-paced environment Robust engineering problem definition Analytical validation and sanity checking Technical documentation Peer review and quality assessment Clear communication of assumptions and engineering decisions Preferred Qualifications: MSc in Aerospace Engineering Prior experience working as Product Owner Offer Details Remote Full-time. 40 hours per week with at least 4 hours PST overlap Location: Bangladesh, Brazil, Colombia, Egypt, Ghana, India, Indonesia, Turkey, Vietnam Engagement Length: 5 weeks Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 8/28/2026
Role Overview We are looking for highly creative and experienced Graphic / Media Designers and Video & Audio Editors to create, edit and evaluate visual and creative content using structured rating frameworks and detailed guidelines. You’ll play a key role in bringing ideas to life through compelling visuals, balancing creativity with precision, and collaborating closely with cross-functional teams to deliver impactful artistic solutions. Job Responsibilities Create high-quality illustrations/videos/visual assets based on creative briefs Produce multiple design/audio/video variations while maintaining consistency in style, composition, and quality Screen and stream capture, creator video management Image editing, photo retouching, compositing, image authoring Digital painting, brush engines, animation, pressure input Non-linear video editing, multi-track timelines, effects, proxy editing Record/Edit audio/video Document prompts, workflows, and asset metadata in a clear and structured manner Collaborate with cross-functional teams and incorporate feedback effectively and efficiently Required Qualification At least 5+ years of professional experience as Graphic / Media Designer, Illustrator, Visual Artist, Video Editor, Audio Engineer, or similar Proficient in 2+ of the listed tools: GIMP, Krita, Inkscape, Kdenlive, Ardour, Audacity, OBS, Jellyfin Strong portfolio demonstrating versatility Analytical mindset with strong problem-solving skills Ability to work independently in a fast-paced environment Strong attention to detail with a commitment to quality and consistency Excellent written and verbal communication skills Strong time management skills with the ability to meet deadlines in a fast-paced environment Preferred Qualifications MSc in Fine Arts and Illustration: or Filmmaking and Production, or related field Prior experience working as Product Owner Offer Details Remote Full-time. 40 hours per week with at least 4 hours PST overlap Location: Bangladesh, Brazil, Colombia, Egypt, Ghana, India, Indonesia, Turkey, Vietnam Engagement Length: 5 weeks Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 8/28/2026
We are building an evaluation benchmark for frontier AI browsing agents. Your job is to design research problems that a state-of-the-art AI cannot solve, even with full web access and multiple attempts. This is not a subject matter expert role, nor is it a content-writing role. It is investigative research. You will start from a verifiable fact, work backwards to construct a question that makes that fact extremely hard to locate, and then prove your work with a complete, auditable evidence trail. What You Will Produce A natural-language research question with a short, stable, objectively verifiable answer Clues, each independently checkable, spanning multiple fact types - including dates, people, places, organisations, works, events, records, and quantities - with specific constraints A validation record showing the obvious searches you ran and the results they returned Minimum Qualifications Demonstrated open-web research ability, including locating primary records and navigating government and institutional databases, archives, registries, and PDF documents Precision with sourcing. You cite exact pages, tables, and sections—not just homepages Comfort researching unfamiliar subjects from scratch Advanced written English High tolerance for structured documentation. The evidence trail is the majority of the work Experience with LLM evaluation, red-teaming, or benchmark construction Experience in one or more of the following domains Reference librarianship, archival research, or special collections Investigative journalism or professional fact-checking OSINT, due diligence, KYC, or investigative research Patent, prior-art, or legal-discovery search Genealogy and records research Competitive quizzing or puzzle-hunt construction Nice to Have Experience with LLM evaluation, red-teaming, or benchmark construction Familiarity with JSON and structured data delivery formats Offer Details Remote Full-time. 40 hours per week with at least 4 hours PST overlap Location: Bangladesh, Brazil, Colombia, Egypt, Ghana, India, Indonesia, Turkey, Vietnam, USA Engagement Length: 8 weeks Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 8/26/2026
We are looking for an experienced Data Scientist to join a Consumer Data Science team focused on using data, experimentation, and advanced modeling to drive product growth and improve the consumer experience. In this role, you will partner closely with Product Managers, Engineers, and Data Scientists across areas such as Feeds, Search, Growth, and Content Ecosystems. You will analyze large-scale datasets, define and monitor key product metrics, uncover strategic insights, and develop data science and machine learning solutions that influence product direction. You will also build scalable data assets, including new tables, self-service tools, and reporting dashboards, enabling product and engineering teams to independently explore data and understand changes in topline metrics. Your work will span experimentation, causal inference, anomaly detection, prediction, pattern recognition, and statistical modeling, with a strong emphasis on translating complex analyses into clear recommendations for stakeholders. Required Skills 5+ years of experience in Data Science, Machine Learning, or a related field. Strong proficiency in Python or R. Strong SQL skills and experience working with relational databases. Strong understanding of statistical modeling, machine learning algorithms, causal inference, and experimental design. Hands-on experience designing, analyzing, and interpreting experiments and product metrics. Experience analyzing and processing large-scale datasets. Experience with distributed data technologies such as Spark, Hadoop, or Hive. Experience with ML libraries and frameworks such as scikit-learn, TensorFlow, or PyTorch. Experience developing dashboards, reporting solutions, and reusable data assets. Ability to identify actionable insights and effectively communicate findings and recommendations to technical and non-technical stakeholders. BS, MS, or PhD in Computer Science, Statistics, Mathematics, or a related quantitative field. Bonus Skills Experience working with consumer-facing technology products. Advanced experience with causal inference and A/B testing. Experience with BigQuery. Experience developing machine learning solutions for anomaly detection, prediction, or pattern recognition. Experience partnering directly with Product and Engineering teams on product strategy and decision-making. Offer Details Full-time dedication (40 hours/week). 5-hour overlap with PST. Location: United States Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 8/20/2026
We are looking for a UX Researcher to join the user experience research team of a large consumer technology platform. This role will focus on rapid, usability-driven research to help product and design teams better understand user behavior, needs, and experiences and translate those findings into meaningful product improvements. You will lead tightly scoped research projects, often running in one-week cycles, and own the research process from initial stakeholder conversations and planning through execution, analysis, and delivery of actionable insights. Your work will include moderated and unmoderated usability testing, user interviews, concept testing, and heuristic evaluations. You will collaborate closely with product managers, product directors, designers, engineers, product marketing, and other researchers. While embedded within the research organization, you will support different consumer product teams, requiring the ability to quickly understand new problem spaces, prioritize research needs, and deliver high-quality insights in a fast-moving environment. Required Skills 3+ years of professional experience conducting user research in a product-focused environment Strong interviewing skills, with experience conducting both user and stakeholder interviews Hands-on experience designing and executing usability testing, including moderated and unmoderated studies Experience with qualitative research methods such as user interviews, concept testing, and usability evaluation Experience conducting heuristic evaluations Ability to independently manage the end-to-end research process, from scoping and research planning through execution, analysis, and presentation of findings Ability to collect, analyze, and synthesize research data into clear and actionable product insights Experience working closely with cross-functional partners, including product managers, designers, engineers, and other stakeholders Ability to appropriately scope and prioritize research projects within short timelines Strong written, verbal, interpersonal, and organizational communication skills Ability to work effectively and independently in ambiguous, fast-changing environments Bachelor’s degree or higher in Psychology, Anthropology, Sociology, Statistics, Behavioral Science, Human-Computer Interaction, Information Systems, Computer Science, or another research-related field, or equivalent relevant experience Bonus Skills Previous experience conducting remote user research Consumer-facing product research experience Previous experience working within a large technology company Master’s degree in a research-related discipline Experience partnering with senior product managers, product directors, and design leaders Understanding of quantitative research methods, including their strengths, limitations, and appropriate use throughout different phases of product development Experience supporting multiple product teams or moving between different research areas Ability to adapt quickly to shifting priorities, schedules, and project requirements Strong portfolio demonstrating end-to-end ownership of usability and product research projects Offer Details Remote Full-time dedication (40 hours/week) Location: USA Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 8/12/2026
About the role We're looking for a Senior Product Security Engineer to partner closely with engineering teams and help build secure products from the ground up. In this role, you'll work alongside product and platform engineers to embed security throughout the Software Development Life Cycle (SDLC), ensuring security is a natural part of development rather than an afterthought. You'll lead architecture reviews, perform hands-on security testing, guide secure development practices, and help create a culture where developers can confidently ship secure software. Required Skills Proven experience in Product Security or Application Security with a demonstrated impact on improving product security. Strong hands-on penetration testing experience for web applications and APIs. Deep understanding of modern web application architecture, including SPAs, APIs, authentication and authorization (OAuth 2.0/OpenID Connect), session management, and common attack vectors such as the OWASP Top 10. Experience conducting security architecture reviews and threat modeling. Strong knowledge of secure software development practices and integrating security throughout the SDLC. Experience performing secure code reviews and establishing secure coding standards. Hands-on experience operating and tuning SAST, DAST, and dependency/supply chain security scanning tools. Ability to triage security findings and translate them into actionable remediation plans. Excellent communication skills with the ability to collaborate effectively with software engineers and explain security risks in practical terms. Experience supporting security compliance initiatives such as SOC 2, ISO 27001, or similar frameworks. Bonus Skills Experience with open-source security tools such as OWASP ZAP, Burp Suite Community Edition, Semgrep, Trivy, Grype, or Nuclei. Offensive security certifications such as OSCP. Cloud security experience across AWS, Azure, or Google Cloud Platform. Experience securing containerized environments and Kubernetes. Background working in enterprise or highly regulated environments. Experience improving developer security education and fostering secure engineering practices. Offer Details Remote Full-time dedication (40 hours/week) Location: Brazil only Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 7/10/2026
Required Skills 7+ years of professional software engineering experience in a fast-paced technology environment. Strong proficiency in Python and frameworks such as Django, FastAPI, or Flask. Strong proficiency in JavaScript/TypeScript and modern frameworks such as Node.js and React (or equivalent frameworks like Vue or Angular). Experience designing, developing, and maintaining full-stack applications. Strong knowledge of RESTful API development and GraphQL. Experience with relational databases (PostgreSQL, MySQL) and NoSQL databases. Hands-on experience with cloud platforms such as AWS, GCP, or Azure. Experience building and maintaining CI/CD pipelines. Familiarity with Docker and Kubernetes. Experience with Git and GitHub/GitLab. Strong understanding of software architecture, system design, and engineering best practices. Experience writing clean, well-tested, and maintainable code. Ability to troubleshoot production issues, optimize application performance, and resolve security vulnerabilities. Experience mentoring engineers and participating in code reviews. Excellent written and verbal English communication skills. Bonus Skills Experience integrating CRM platforms, marketing platforms, or other third-party business systems through APIs. Experience designing scalable integration architectures across multiple business applications. Exposure to Go-to-Market systems and business operations tooling. Offer Details Remote Full-time dedication (40 hours/week) Location: India, Bangladesh, Philippines, Singapore, Mexico Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 7/9/2026
We are looking for an experienced Senior Salesforce Administrator to join a fast-paced Go-to-Market Systems team supporting Sales and Operations. In this role, you will serve as the first point of contact for Salesforce users, ensuring the platform remains reliable, accessible, and efficient for day-to-day business operations. You will be responsible for resolving user issues, managing access, performing light system configuration, maintaining data quality, and collaborating with Product and Engineering teams to escalate and troubleshoot more complex platform challenges. Required Skills 5–7 years of hands-on experience as a Salesforce Administrator. Strong knowledge of Salesforce administration, including user provisioning, de-provisioning, and access management. Experience managing roles, profiles, permission sets, sharing rules, and Salesforce security models. Solid understanding of Salesforce standard objects, including Leads, Contacts, Accounts, and Opportunities. Experience troubleshooting user access, visibility issues, and standard Salesforce errors. Ability to perform routine Salesforce configuration, including fields, page layouts, list views, reports, and other declarative features. Experience supporting data quality initiatives through imports, exports, and data maintenance. Familiarity with monitoring Salesforce Flows, Validation Rules, and other standard automations. Strong analytical and problem-solving skills with the ability to identify root causes before escalating issues. Excellent written and verbal communication skills with a customer-focused mindset. Experience documenting processes, troubleshooting guides, and standard operating procedures. Bonus Skills Salesforce Certified Administrator (ADM-201). Experience supporting Go-to-Market, Sales, or Revenue Operations teams. Familiarity with ticketing and support queue management. Experience collaborating with Product and Engineering teams in cross-functional environments. Knowledge of Salesforce best practices for system administration and governance. Offer Details Remote Full-time dedication (40 hours/week) Location: India, Bangladesh, Philippines, Singapore, Mexico Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 7/9/2026
We are looking for a creative and detail-oriented Senior Illustrator/Sketcher/Cartoonist with strong visual skills and the ability to work across diverse artistic sketching styles. The role requires both high-quality illustration work and clear, professional communication. Key Responsibilities Create stylized illustrations based on briefs while maintaining composition and visual consistency Adapt and replicate multiple illustration styles and produce required variants with accuracy Ensure quality in lighting, texture, and overall output Document prompts, style choices, and metadata clearly Collaborate effectively with teams and incorporate feedback Ensure final outputs are polished, production-ready, and aligned with project goals Style Expertise Sketch & Drawing Cinematic & Film Noir Steampunk Surreal / Dreamlike / Psychedelic Cyberpunk & Sci-Fi Art Deco & Art Nouveau Vintage & Retro Glitch & Vaporwave Tools Adobe Photoshop Procreate Adobe Illustrator (for vector illustration, if required) Affinity Suite Clip Studio Paint Requirements Minimum 3+ years of professional experience as an Illustrator/Sketcher/Cartoonist or in a similar role Strong portfolio showcasing versatility across multiple styles Strong fundamentals in composition, color, and lighting Ability to paint/render scenes, characters, or environments Ability to replicate multiple art styles accurately Good written and verbal communication skills High attention to detail and time management Location: Argentina, Mexico, Bangladesh, Sri Lanka, Turkey, Brazil, Egypt, Ghana, India. Commitments Required: Full-time (8 hours/day) 40 hours per week. Overlap 4 hours with PST Engagement type: Contractor Duration of contract : 4 Weeks Important compensation context:This is a short-term contractor engagement (4 weeks, full-time) focused on AI training data — specifically, creating illustrations and sketches that help train and evaluate AI models. Tasks are self-contained and range from 10 to 30 minutes each, following clear creative briefs. You'll work independently at your own pace within a structured task queue. Because compensation is tied to a pay-per-task model (not a traditional hourly salary), the effective hourly rate will vary based on task type and volume. This structure is best suited for candidates who are comfortable with task-based, gig-style work rather than a fixed monthly income. Please share your minimum acceptable hourly rate so we can assess alignment before moving forward. Candidates with rate expectations significantly above the market range for AI annotation and data labeling work may not be a strong fit for this particular engagement. Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 6/25/2026
Role Overview We are hiring for a PAA-based aerospace project, seeking candidates with strong hands-on experience in CAD modeling and simulation tools. This role is ideal for experienced aerospace engineering professionals with deep expertise in aerospace design, CAD engineering, and simulation workflows. Key Responsibilities Develop CAD models and perform simulation/analysis for aerospace components Validate designs using industry-standard tools (e.g., CATIA, SolidWorks, ANSYS, MATLAB) Translate technical requirements into simulation models and workflows Optimize designs based on simulation results Collaborate with cross-functional engineering teams Document methodologies, processes, and outcomes Key Requirements: Proficiency in CAD and simulation tools Candidates must have a minimum of 8+ years of experience working in CAD engineering and simulation-focused organizations Strong aerospace/mechanical engineering fundamentals Hands-on experience with modeling, simulation, or design projects Analytical mindset with strong problem-solving skills Ability to work independently in a fast-paced environment Preferred Qualifications: Strong proficiency in CAD engineering and simulation tools Solid aerospace/mechanical engineering fundamentals Hands-on experience in aerospace modeling, simulation, and design projects Strong analytical and problem-solving skills Ability to work independently in a fast-paced engineering environment Prior experience working with leading aerospace or engineering companies is highly preferred Offer Details: Employment Type: Contractor assignment Duration of contract : 8 weeks; [expected start date is next week] Commitment Required: Full-time. 40 hours per week with an overlap of 4 hours with PST. Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 6/1/2026
About Mavila Consulting Mavila Consulting is a specialized recruitment partner connecting exceptional talent from around the world with world-class companies (mostly from the USA). Over the past years, we’ve helped 600+ professionals land remote roles with global teams across engineering, product, data, and operations. About this Talent Database This form is not for a specific opening. It feeds our internal Talent Database, which we use to quickly match candidates to new roles as they open. When clients launch a search, we start here reviewing profiles, skills, and preferences to invite the best-fit professionals to the next steps. What We’re Looking For Tell us about your skills, experience, languages, location and compensation expectations. The more complete your information, the faster we can match you when opportunities arise. What Happens Next After you submit, your profile becomes searchable by our team. As soon as a relevant role appears, we’ll reach out with the job details and guide you through the process. If this sounds good, please complete the form below, we’re excited to learn more about you! IMPORTANT: We’re not able to provide individual feedback or reply to messages sent through this form. Your details go into our Talent Database, and we’ll reach out only if there’s a strong match. We handle ~30 roles per month, so if your skills align, you may hear from us. Thanks for understanding!
Posted 5/14/2026
Role Overview Evaluate and compare the quality of responses from multiple AI chatbots across real-world small business use cases. Responsibilities Create realistic business-related prompts based on defined user goals Interact with multiple AI chatbots (max. 5 turns per conversation) Assess response quality across clarity, usefulness, and accuracy Provide structured feedback and comparative evaluations Submit conversation transcripts and evaluation results Requirements Business owner or strong understanding of small business operations Fluent in Portuguese (Brazilian) Strong analytical and critical thinking skills Ability to follow structured evaluation guidelines Comfortable interacting with AI tools 2+ years of experience Intermediate to Advanced English What you'll work on: Create engaging visual content for marketing Help answer and evaluate situations related to day-to-day operations and customer interactions Conduct market research and contribute ideas in your area of expertise Work with data to support analysis and financial planning Review and evaluate AI-generated responses for small business use cases Use tools and input files such as spreadsheets, PDFs, and images as part of your workflow Offer Details: Project-based with defined number of evaluation tasks Each task includes multi-chatbot comparison and final assessment Duration: 16 weeks Full-time. 40 hours per week with at least 4 hours PST overlap Brazil only (portuguese speakers) Observations: Marketing content creation (visual) At least 30% of conversations user should supply their own business logo or product images Generating or manipulating visual media such as logo, campaigns, flyers, designs, professional product catalog and artwork. Users want to bring visual ideas to life or modify existing visuals. Daily Operations & Customer Management At least 50% of the conversations users should supply file inputs. Coordinating daily workflows, inventory logistics, team schedules, and automating CRM tasks. Users want to eliminate tedious manual data entry and organize their day-to-day business operations efficiently without relying on specialized software. Market Intelligence & Ideation Researching competitor landscapes and target audience behaviors to define Ideal Customer Profiles (ICPs) and pinpoint market saturation. Users want to understand their customers' deep-seated needs and build strategic, SEO-driven roadmaps to launch, grow, or monetize a business. Data analysis & financial planning At least 80% of the conversations users should supply file inputs. Handling budgeting, cash flow tracking, bookkeeping, and streamlined pricing and quoting workflows. Users want to manage their financial runway, understand real-time profitability, and generate quick, accurate estimates to win local business without relying on a dedicated accountant. The business type really doesn’t matter. It can be stores, restaurants, gyms, aesthetic clinics, beauty salons, and so on. Attention: This is a short-term project with a duration of 16 weeks. ATENÇÃO! Esta é uma posição voltada para empreendedores de pequenos negócios, como salão de beleza, loja, restaurante, academia, clínica, entre outros. O CV será o principal documento analisado pelo cliente. Por isso, é importante que o seu negócio esteja incluído como experiência profissional, com descrição das atividades, responsabilidades e competências desenvolvidas. O CV deve estar atualizado, com o negócio já incluído, e obrigatoriamente em inglês. Important note: Due to the high volume of applications, we are unable to provide individual feedback at this stage. Only candidates shortlisted for the next steps will be contacted by our team.
Posted 5/5/2026
Browse hundreds of curated positions from companies hiring in Latin America.