Research | WRITER
Advancing AI for the enterprise
At WRITER, we have one goal: to build scalable, reliable,
and transparent AI technology for the enterprise.
Our approach
Our approach is different: we believe that building large language models (LLMs) informed by enterprise requirements leads to AI systems that are more reliable, more controllable, and more transparent. When you ground cutting-edge AI innovation in real-life needs, it yields solutions that solve problems people actually face.
Our team
Our globally distributed team of AI/ML researchers and engineers has a five-year track record of groundbreaking research and development across language models, retrieval systems, and evaluations.
Our results
Our results demonstrate that when AI research starts with real needs, it leads to:
- Prioritizing capabilities that map to tangible outcomes
- Balanced focus between sophistication and practicality
- Better evaluation metrics to understand real-world performance
- Earlier identification of potential risks and failures
Research highlights
Research pillars
Enterprise-optimized models
Focus on developing more scalable, reliable, and transparent models specifically engineered for enterprise requirements
Practical evaluations
Development of model evaluation methodology that reflects real-world scenarios and risks
Domain-specific specialization
Research into applying AI systems in high-stakes industries
Retrieval & knowledge integration
Work on next-generation retrieval systems that safely and reliably connect language models with enterprise data
All Practical evaluations Enterprise-optimized models Domain-specific specialization Retrieval
How personalized context quietly degrades AI accuracy: a deeper look
Practical evaluations
Jul 26, 2026
Accurate Failure Prediction in Agents Does Not Imply Effective Failure Prevention
Practical evaluations
Feb 3, 2026
Towards Outcome-Oriented, Task-Agnostic Evaluation of AI Agents
Practical evaluations
Nov 11, 2025
Palmyra-mini: Small models, big throughput, powerful reasoning
Enterprise-optimized models
Sep 19, 2025
Reflect, retry, reward: Self-improving LLMs via reinforcement learning
Practical evaluations
Jun 12, 2025
Palmyra X5: The end of context constraints
Enterprise-optimized models
Apr 28, 2025
Expecting the unexpected: FailSafeQA Benchmark
Practical evaluations,
Domain-specific specialization
Feb 10, 2025
Palmyra Creative: Unlocking creativity with AI
Domain-specific specialization
Dec 17, 2024
Introducing Self-evolving models
Enterprise-optimized models
Nov 20, 2024
Palmyra X4: Introducing actions
Enterprise-optimized models
Oct 9, 2024
Writing in the Margins
Enterprise-optimized models
Aug 27, 2024
Comparative analysis of retrieval systems in the real-world
Retrieval
May 3, 2024
OmniACT: A benchmark for enabling multimodal generalist autonomous agents
Practical evaluations
Feb 27, 2024
Fusion-in-decoder: achieving state-of-the-art open-domain QA performance
Enterprise-optimized models
Sep 13, 2023
Becoming self-instruct: Introducing early stopping criteria for minimal instruct tuning
Enterprise-optimized models
Jul 5, 2023
Palmyra Med: Instruction-based fine-tuning of LLMs enhancing medical domain performance
Domain-specific specialization
Jul 3, 2023
Palmyra Fin
Domain-specific specialization
Jul 3, 2023
Grammatical error correction: a survey of the state of the art
Enterprise-optimized models
Apr 29, 2023