Turkey's Homegrown, Original AI Technology
Olgu AI is a Turkey-based AI company building large language models built specifically for Turkish.
Our story
Existing large language models usually treat Turkish as a second-class language — they perform impressively in English and a handful of other major languages, but often produce weaker, less natural Turkish. Olgu AI was born from the belief that our mother tongue, spoken by over 84 million people, deserves to reach its rightful competence in the age of AI, and we're building a from-scratch Turkish architecture to close that gap.
Our mission
From individuals to institutions and the public sector, we build models that genuinely understand Turkish culture and linguistic nuance. Thanks to our from-scratch, Turkish-specific tokenizer, these models run more efficiently — meaning lower cost for institutions and more accessible pricing for individuals — while your data always stays inside the country.
Who we build for
Olgu AI aims to create value across a broad spectrum of users.
Individuals
An AI that's there for your daily work, classes, homework, or everyday conversations. Our models understand Turkish nuance and culture with a natural fluency no other model matches, and thanks to our efficient tokenizer architecture, offer far more accessible pricing than global alternatives.
Enterprises & Public Sector
We offer two paths for private-sector companies and public institutions alike: purpose-built, ultra-efficient small language models (SLMs) tailored to your domain — with your data staying inside your organization — or direct API access to our model family for fast integration.
National Security
A fully domestic, independent, and auditable AI infrastructure for critical institutions and the defense industry. A strategic technology base that removes foreign dependency.
Olgu AI by the numbers
The core metrics we measured and validated while building our first model.
54.7%
More efficient tokenization
Our from-scratch, Turkish-specific tokenizer represents the same Turkish text 54.7% more efficiently than LLaMA-based open-source models.
2.21×
Fewer tokens per word
That translates into lower operating cost and a wider effective context window on the same hardware. In practice, for the same budget, we can process roughly 121% more conversation history than a similarly sized foreign model.
1B
Parameters in our first model
Our first model, trained entirely from scratch — not just to validate tokenizer efficiency, but as a full end-to-end run of our entire approach, from data collection through training, and the foundation for the larger models to come.
100%
Open, Honest Benchmarking
When we compare our model against other Turkish language models, we publish where we fall short, not just where we lead.
Why a national language model?
Most large language models are optimized for English; languages like Turkish are usually an afterthought. We're building the alternative.
Data Sovereignty & Security: Data infrastructure that never leaves the country — KVKK-compliant, processed on domestic servers.
Cost & Context Advantage: 54.7% lower processing cost thanks to our Turkish-specific tokenizer.
Native-Language Naturalness: Unlike English-centric models, full alignment with Turkish grammar and cultural nuance.
Industry-Specific Mini Models: For any local business or public institution, we design super-efficient mini Turkish models trained on that domain's own data — delivering far stronger performance at a fraction of the cost of foreign or fine-tuned alternatives.
Removing Foreign Dependency: Your AI infrastructure doesn't depend on another country keeping its APIs open or publishing its models publicly — no nation can build a critical industry on that kind of dependency.
Potential for Defense Industry Integration: Foreign-sourced AI technology can't be effectively used in the defense industry for security and sovereignty reasons; our fully domestic models are positioned to integrate into that space.
2.21×
Fewer Tokens Per Word
121% More Conversation HistoryOur approach
Rather than adapting an existing model to Turkish after the fact, we follow a Turkish-first process from the ground up.
Data collection & cleaning
We carefully assemble Turkish text data, using MinHash-based near-duplicate detection with sample-based verification to safeguard data quality.
Turkish-first tokenizer
Our from-scratch tokenizer represents Turkish text 54.7% more efficiently than LLaMA-based open-source models — 2.21× fewer tokens per word.
Training from scratch
Rather than adapting an existing model to Turkish after the fact, we train from the ground up around Turkish's own linguistic structure.
Our Contribution to Turkish NLP
We're not just building models — we aim to give back what we learn to the field of Turkish natural language processing.
Technical Reports & Papers
From tokenizer design to training methodology, we'll be regularly publishing our methods and findings.
Open Experiments & Learnings
We share not just what worked, but what we tried and what didn't — contributing to the field's collective progress.
Collaboration with Academia & Community
We're open to working with Turkish NLP researchers, universities, and the open-source community.
Our values
Transparency
We share openly what we do and how we do it, describing the model's capabilities and limits honestly. As part of that transparency, we regularly publish what we learn through technical reports and research to contribute to the field of Turkish NLP.
Responsibility
We hold user data and KVKK compliance to the highest standard — you stay in control of your data. Everything we build operates within existing laws and regulations, for the benefit of society, benefiting people, and without violating anyone's rights.
Data sovereignty
Critical and public-sector institutions' data is processed and stored within Turkey's borders, without being sent abroad.
Accessibility
AI should be easily accessible to everyone, with a simple experience that requires no technical knowledge.
Contribution to National Security
AI technology is becoming an increasingly critical field for national security. As a fully domestic, auditable technology base, we aim to be positioned to contribute to defense and national-security institutions over the long term.
Fair, Local Pricing
We make sure people can access highly effective models built exactly for their needs at fair, locally affordable prices — without being overcharged by international companies — a strategy that enables an independent private sector built on domestic technology.
Our roadmap
Our journey to build Turkey's own AI ecosystem.
First Model: Field Validation
1B Parameter ClassWe built our first 1-billion-parameter model from scratch and validated our Turkish-specific tokenizer's efficiency advantage in the field — the architectural foundation for the larger-scale models to come.
Scaling Up
32B – 64B Parameter ClassGrowing our model to a frontier-adjacent 32B–64B parameter class, targeting strong performance on complex reasoning and real-world enterprise scenarios.
National Scale & Core Infrastructure
100B+ Parameter ClassBecoming Turkey's de facto domestic AI provider — an independent technology base supporting local industry, public-sector projects, startups, and the everyday productivity of every user in Turkey.
