Appen’s cover photo
Appen

Appen

IT Services and IT Consulting

Kirkland, Washington 1,074,600 followers

Appen is your trusted data partner, powering cutting-edge AI applications for the world's most innovative companies.

About us

Appen has been a leader in AI training data for over 25 years, providing high-quality, diverse datasets that power the world's leading AI models. Our end-to-end platform, deep expertise, and scalable human-in-the-loop services enable AI innovators to build and optimize cutting-edge models. We specialize in creating bespoke, human-generated data to train, fine-tune, and evaluate AI models across multiple domains, including generative AI, large language models (LLMs), computer vision, speech recognition, and more. Our solutions support critical AI functions such as supervised fine-tuning, reinforcement learning with human feedback (RLHF), model evaluation, and bias mitigation. Our advanced AI-assisted data annotation platform, combined with a global crowd of more than 1M contributors in over 200 countries, ensures the delivery of accurate and diverse datasets. Our commitment to quality, scalability, and ethical AI practices makes Appen a trusted partner for enterprises aiming to develop and deploy effective AI solutions. At Appen, we foster a culture of innovation, collaboration, and excellence. We value curiosity, accountability, and a commitment to delivering the highest-quality AI solutions. We support work-life balance with flexible work arrangements and a dynamic, results-driven environment. Employees have access to competitive pay, comprehensive benefits, and opportunities for continuous learning and career growth. Our team works closely with the world’s top technology companies and enterprises, tackling exciting challenges and shaping the future of artificial intelligence.

Industry
IT Services and IT Consulting
Company size
501-1,000 employees
Headquarters
Kirkland, Washington
Type
Public Company
Founded
1996
Specialties
Search, Annotation, Evaluation, Personalization, Transcription, Spam Detection, Translation and Localization, Data Collection, training data, artificial intelligence , machine learning, data preparation, model evaluation, datasets, computer vision, natural language processing, LLM, and generative ai

Locations

Employees at Appen

Updates

  • Appen reposted this

    While the frontier labs race towards bigger models and more compute, we believe the next wave of AI will be defined by more efficient architectures. Our Co-Founder & CTO, Alexander Whedon, recently joined Karla Heredia and Sergio Bruccoleri from Appen to discuss why architecture is becoming the defining factor in the next generation of AI systems. A few topics explored: → Why transformer models become increasingly inefficient as context grows. → Why today's AI agents spend too much time managing context instead of reasoning. → How long-horizon memory will change enterprise AI and autonomous agents. → Why independent benchmarking is critical for evaluating long-context models. → How long context allows developers to simplify RAG systems by reducing chunking, retrieval complexity, and manual orchestration. A great conversation on where foundation models are headed next. Watch the full discussion. Link in the comments.

  • View organization page for Appen

    1,074,600 followers

    Healthcare AI may be approaching a bottleneck that better architectures and more compute won’t solve: access to the right data. Clinical AI is moving beyond narrow benchmarks into multimodal systems that need to reason across imaging, clinical notes, biosignals and physician-patient dialogue. But the data required to train and evaluate those systems is unusually difficult to scale. It has to be privacy-safe and traceable. It often requires credentialed medical experts to annotate and adjudicate. And for many emerging use cases, public or synthetic datasets don’t capture the clinical variability models encounter in deployment. The problem becomes even sharper with clinical agents. Recent research shows how dramatically performance can change when models move from static medical questions to sequential, tool-using doctor-patient interactions. That means evaluation itself needs richer clinical data: longitudinal dialogue, multiple specialties and languages, realistic tool use, and expert validation. As healthcare models become more capable, the limiting resource may increasingly be high-quality, expert-validated clinical data rather than model capacity. Our latest Appen Insights looks at the coming medical data crunch, and why data access, provenance and clinical expertise could become a competitive moat for healthcare AI. Read the full blog by Sergio Bruccoleri, VP of Delivery at Appen, in the comments. #HealthcareAI #AIResearch

  • View organization page for Appen

    1,074,600 followers

    If an LLM calls a solver and gets the right answer, did the model reason? Or did the system reason? That distinction sits at the center of our latest episode of The Data Layer with Dan Roth, Chief AI Scientist at Oracle and Distinguished Professor at the University of Pennsylvania, whose research spans machine learning, reasoning, language understanding, and neurosymbolic AI. Dan makes a sharp argument: today’s language models are inductive learning machines. There are classes of reasoning problems they cannot reliably solve on their own. But systems built around them can, by knowing when to retrieve information, invoke specialized solvers, use tools, and verify the result. And that creates a second problem: how do we evaluate the system rather than just the model? Dan joins Appen’s Jeanine Sinanan-Singh, Director of GenAI Research and Brian Jenkins, VP of Sales & Marketing at Appen, to get into what that means for agentic AI: why retrieval is still an open problem, why benchmark performance can hide real-world failures, why evaluation needs to account for messy production environments, and why runtime monitoring may become just as important as offline evals. One idea from the conversation that stuck with us: the harder a failure is for a user to see, the more important it becomes for the system itself to detect it. For anyone working on agents, reasoning, retrieval, or evaluation, this is a good one. Full conversation in comments. #AIResearch #AgenticAI #AIEvaluation #LLM 

    • No alternative text description for this image
  • View organization page for Appen

    1,074,600 followers

    What will differentiate the next generation of AI models? At SlatorCon London 2026, Appen's Sergio Bruccoleri, Appen's VP of Delivery, joined Anna Wyndham (Slator) to discuss the evolving Data for AI market. The discussion touched on several themes shaping the industry: • The shift toward high-value, highly subjective use cases such as coding, STEM, healthcare, legal, and HR. • Growing demand for domain experts rather than generalist contributors. • The importance of expertise across languages, cultures, and modalities. • The challenge of capturing tacit knowledge developed through experience. Read Slator's recap of the discussion. Link in the comments. #AI #DataForAI #LLMs #MachineLearning #HumanInTheLoop

    • No alternative text description for this image
  • View organization page for Appen

    1,074,600 followers

    A benchmark pass can still be an evaluation failure. Poolside’s Laguna S 2.1 highlights a growing problem for agent researchers: reward hacking. An agent can reach the correct outcome without demonstrating the capability the benchmark was designed to measure. Poolside’s evaluation methodology looks beyond pass rates. It publishes full trajectories and combines human-labeled calibration, adversarial judging and expert review to identify suspicious behavior. Appen worked with Poolside on aspects of the reward hacking detection used during the evaluation process. In our latest article, Jeanine Sinanan-Singh, Director of GenAI Research at Appen, examines why agent evaluations must measure the process, not only the outcome, and why benchmark results need a stronger audit trail. Link in comments.   #AgentEvaluation #RewardHacking

    • No alternative text description for this image
  • View organization page for Appen

    1,074,600 followers

    Great to see the Poolside team release Laguna S 2.1. As AI agents tackle longer, more complex workflows, efficient long-context reasoning and persistent agentic capabilities are becoming increasingly important. Congratulations to the Poolside team on the release. Looking forward to seeing how the community builds with it. #AI #AgenticAI #OpenSourceAI #LLMs

    View organization page for Poolside

    30,946 followers

    Today we’re releasing Laguna S 2.1, our new open-weight model for agentic coding and long-horizon work. Laguna S 2.1 is a 118B-parameter Mixture-of-Experts model with 8B active parameters per token, up to 1M tokens of context, and thinking and no-thinking modes. It is capable enough to compete with models several times its size, yet small enough to run locally on a single NVIDIA DGX Spark. What sets Laguna S 2.1 apart is its persistence. Across long-horizon coding and research tasks, it holds onto a goal, uses tools, checks its work, recovers when an approach fails, and continues making progress for hours with little or no intervention. That persistence comes with a practical balance of cost, speed, and ownership. Laguna S 2.1 weight class is designed to take on real, long-running agentic work at a cost and speed that make it practical to run often and at scale. Because it activates only 8B parameters per token, long agent runs and reinforcement learning loops are faster and less expensive than they would be with much larger models. We’re releasing Laguna S 2.1 under OpenMDW-1.1 with checkpoints in BF16, FP8, INT4, and NVFP4, alongside official GGUF and MLX quantizations. The weights are available today on Hugging Face. Run it through pool, vLLM, SGLang, Ollama, llama.cpp, ZML, MLX, or NVIDIA TensorRT-LLM, or access it through OpenRouter and the Poolside API.

    • No alternative text description for this image
  • View organization page for Appen

    1,074,600 followers

    A correct answer doesn't always mean correct reasoning. As LLMs become part of enterprise workflows, evaluating outputs alone is no longer enough. A model can generate the right answer while relying on brittle shortcuts, superficial pattern matching, or reasoning that doesn't hold up when the prompt, policy, or context changes. That's why reasoning quality is becoming an increasingly important part of AI evaluation. Organizations need to understand not only what a model produces but how it reaches its conclusions. In Appen's latest article with The AI Journal, Jeanine Sinanan-Singh explores why reasoning-aware evaluation matters and how approaches such as adversarial testing, process-based evaluation, and stronger reasoning datasets can help build more reliable AI systems. As AI systems become more capable, should reasoning quality become a standard evaluation metric alongside accuracy? Read the full article in the comments. #AIResearch #LLMs #GenerativeAI

    • No alternative text description for this image
  • View organization page for Appen

    1,074,600 followers

    OpenAI's GPT-Red is a big step for automated red-teaming but it doesn't close the safety evaluation gap. It shifts it. GPT-Red is trained via self-play against a population of defender models, each environment built around an explicit threat model. OpenAI reports using its attacks to train GPT-5.6 Sol, with 6x fewer failures on its hardest direct prompt-injection benchmark and 97%+ accuracy on several indirect prompt-injection benchmarks. The transfer results are genuinely notable: in an internal replication of a published indirect prompt-injection arena, GPT-Red succeeded in 84% of scenarios against GPT-5.1 compared to 13% for human red-teamers. It also transferred attacks from simulation to a real deployed agent, and outperformed a prompted baseline attacking a Codex CLI agent. But our team's read: these are strong robustness results within specific, largely internal benchmarks, not evidence that safety generalizes uniformly across threat classes. Appen's research (Same Model, Different Weakness, 363 adversarial scenarios across US English and Mexican Spanish, 52,272 human harm ratings) found something OpenAI's disclosure doesn't yet address: language and modality interact. Role-play attacks got weaker in Spanish; visual attacks got stronger. Model safety rankings didn't hold across languages. That's the core issue with any single-lab red-teaming result, however impressive: benchmark saturation (>97%) tells you a suite has lost discriminatory power for one model, not that the underlying threat is solved. Our full breakdown covers what GPT-Red's evaluations do and don't establish and what independent evaluation design needs to look like next. Link in the comments. 

    • No alternative text description for this image

Similar pages

Browse jobs