About us
Olgu AI is a Turkey-based AI company building large language models built specifically for Turkish.

Our story
The idea for Olgu AI came from watching existing large language models treat Turkish as an afterthought. These models perform impressively in English and a handful of other major languages, but often produce weaker, less natural, sometimes outright incorrect Turkish. That gap is what Olgu set out to close — built on the belief that a language spoken by over 84 million people shouldn't be left behind in the age of AI.
Our mission
Olgu's mission is to be the strong, homegrown AI company Turkey needs — for individuals in daily life, for institutions as trustworthy infrastructure, and over the long term as a domestic technology base in national security as well. We build models that speak Turkish at a native level and genuinely understand Turkish culture and the intricacies of the language. Thanks to our from-scratch, Turkish-specific tokenizer, our models don't just process Turkish more efficiently — they think in Turkish, an efficiency that means lower cost for institutions and meaningfully more affordable pricing for individual users compared to foreign platforms. At the same time, we aim to offer trustworthy AI infrastructure that keeps data inside the country instead of sending it abroad.
Who we build for
Olgu AI aims to create value across a broad spectrum of users.
Individuals
An AI that's there for your daily work, learning, and everyday conversations — one that genuinely understands Turkish and Turkish culture. The efficiency of our Turkish-specific tokenizer lets us offer meaningfully lower pricing than foreign platforms for the same usage, a benefit that goes straight to the people using it.
Institutions
We aim to give Turkish companies and public institutions a trustworthy AI infrastructure they can adapt to their needs, without their data ever leaving the country.
National security
Over the long term, we see a path to becoming a fully domestic, auditable technology base capable of partnering with defense and national-security-related institutions.
Olgu AI by the numbers
The core metrics we measured and validated while building our first model.
54.7%
More efficient tokenization
Our from-scratch, Turkish-specific tokenizer represents the same Turkish text 54.7% more efficiently than LLaMA-based open-source models.
2.21×
Fewer tokens per word
That translates into lower operating cost and a wider effective context window on the same hardware. In practice, for the same budget, we can process roughly 121% more conversation history than a similarly sized foreign model.
1B
Parameters in our first model
Our first model, trained entirely from scratch — not just to validate tokenizer efficiency, but as a full end-to-end run of our entire approach, from data collection through training, and the foundation for the larger models to come.
Honest
Our benchmarking principle
When we compare our model against other Turkish language models, we publish where we fall short, not just where we lead.
Why a national language model?
Most large language models are trained overwhelmingly on English and a handful of other major languages. Languages like Turkish are usually an afterthought in these models — both in linguistic naturalness and in processing efficiency. We're building a model that understands and produces Turkish at a native level, and that — thanks to our Turkish-specific tokenizer — can process more Turkish content on the same hardware, meaning both lower operating cost and stronger context understanding. AI is becoming an increasingly critical technology, from everyday personal use to enterprise operations, public services, and national security; at that scale, foreign dependency carries real long-term strategic risk for any country. Today, many organizations meet their Turkish-language needs by fine-tuning English-centric models — an approach that's both inefficient and dependent on foreign infrastructure. We're building the alternative.
2.21×
fewer tokens per word — the efficiency gain measured with our Turkish-specific tokenizer
Our approach
Rather than adapting an existing model to Turkish after the fact, we follow a Turkish-first process from the ground up.
Data collection & cleaning
We carefully assemble Turkish text data, using MinHash-based near-duplicate detection with sample-based verification to safeguard data quality.
Turkish-first tokenizer
Our from-scratch tokenizer represents Turkish text 54.7% more efficiently than LLaMA-based open-source models — 2.21× fewer tokens per word.
Training from scratch
Rather than adapting an existing model to Turkish after the fact, we train from the ground up around Turkish's own linguistic structure.
Our values
Transparency
We share openly what we do and how we do it, describing the model's capabilities and limits honestly.
Responsibility
We hold user data and KVKK compliance to the highest standard — you stay in control of your data.
Data sovereignty
Critical and public-sector institutions' data is processed and stored within Turkey's borders, without being sent abroad.
Accessibility
AI should be easily accessible to everyone, with a simple experience that requires no technical knowledge.
Our roadmap
First model: proof of concept
We built our first 1-billion-parameter model from scratch and validated the efficiency advantage of our Turkish-specific tokenizer — the foundation for the larger models to come.
Scaling up
Growing our model to several billion parameters, targeting genuinely competitive performance in real-world use cases.
National scale
Building a large-scale Turkish AI model that meets the nation's needs and is competitive on an international level — a genuine alternative for individuals, organizations, and the public sector.
