olgu.aiOlgu AI
About us

About us

Olgu AI is a Turkey-based AI company building large language models built specifically for Turkish.

Our story

The idea for Olgu AI came from watching existing large language models treat Turkish as an afterthought. These models perform impressively in English and a handful of other major languages, but often produce weaker, less natural, sometimes outright incorrect Turkish. That gap is what Olgu set out to close — built on the belief that a language spoken by over 84 million people shouldn't be left behind in the age of AI.

Our mission

Olgu's mission is to be the strong, homegrown AI company Turkey needs — for individuals in daily life, for institutions as trustworthy infrastructure, and over the long term as a domestic technology base in national security as well. We build models that speak Turkish at a native level and genuinely understand Turkish culture and the intricacies of the language. Thanks to our from-scratch, Turkish-specific tokenizer, our models don't just process Turkish more efficiently — they think in Turkish, an efficiency that means lower cost for institutions and meaningfully more affordable pricing for individual users compared to foreign platforms. At the same time, we aim to offer trustworthy AI infrastructure that keeps data inside the country instead of sending it abroad.

Who we build for

Olgu AI aims to create value across a broad spectrum of users.

Individuals

An AI that's there for your daily work, learning, and everyday conversations — one that genuinely understands Turkish and Turkish culture. The efficiency of our Turkish-specific tokenizer lets us offer meaningfully lower pricing than foreign platforms for the same usage, a benefit that goes straight to the people using it.

Institutions

We aim to give Turkish companies and public institutions a trustworthy AI infrastructure they can adapt to their needs, without their data ever leaving the country.

National security

Over the long term, we see a path to becoming a fully domestic, auditable technology base capable of partnering with defense and national-security-related institutions.

Olgu AI by the numbers

The core metrics we measured and validated while building our first model.

54.7%

More efficient tokenization

Our from-scratch, Turkish-specific tokenizer represents the same Turkish text 54.7% more efficiently than LLaMA-based open-source models.

2.21×

Fewer tokens per word

That translates into lower operating cost and a wider effective context window on the same hardware. In practice, for the same budget, we can process roughly 121% more conversation history than a similarly sized foreign model.

1B

Parameters in our first model

Our first model, trained entirely from scratch — not just to validate tokenizer efficiency, but as a full end-to-end run of our entire approach, from data collection through training, and the foundation for the larger models to come.

Honest

Our benchmarking principle

When we compare our model against other Turkish language models, we publish where we fall short, not just where we lead.

Why a national language model?

Most large language models are trained overwhelmingly on English and a handful of other major languages. Languages like Turkish are usually an afterthought in these models — both in linguistic naturalness and in processing efficiency. We're building a model that understands and produces Turkish at a native level, and that — thanks to our Turkish-specific tokenizer — can process more Turkish content on the same hardware, meaning both lower operating cost and stronger context understanding. AI is becoming an increasingly critical technology, from everyday personal use to enterprise operations, public services, and national security; at that scale, foreign dependency carries real long-term strategic risk for any country. Today, many organizations meet their Turkish-language needs by fine-tuning English-centric models — an approach that's both inefficient and dependent on foreign infrastructure. We're building the alternative.

2.21×

fewer tokens per word — the efficiency gain measured with our Turkish-specific tokenizer

Our approach

Rather than adapting an existing model to Turkish after the fact, we follow a Turkish-first process from the ground up.

Data collection & cleaning

We carefully assemble Turkish text data, using MinHash-based near-duplicate detection with sample-based verification to safeguard data quality.

Turkish-first tokenizer

Our from-scratch tokenizer represents Turkish text 54.7% more efficiently than LLaMA-based open-source models — 2.21× fewer tokens per word.

Training from scratch

Rather than adapting an existing model to Turkish after the fact, we train from the ground up around Turkish's own linguistic structure.

Our values

Transparency

We share openly what we do and how we do it, describing the model's capabilities and limits honestly.

Responsibility

We hold user data and KVKK compliance to the highest standard — you stay in control of your data.

Data sovereignty

Critical and public-sector institutions' data is processed and stored within Turkey's borders, without being sent abroad.

Accessibility

AI should be easily accessible to everyone, with a simple experience that requires no technical knowledge.

Our roadmap

Now

First model: proof of concept

We built our first 1-billion-parameter model from scratch and validated the efficiency advantage of our Turkish-specific tokenizer — the foundation for the larger models to come.

Near future

Scaling up

Growing our model to several billion parameters, targeting genuinely competitive performance in real-world use cases.

Long term

National scale

Building a large-scale Turkish AI model that meets the nation's needs and is competitive on an international level — a genuine alternative for individuals, organizations, and the public sector.

Join us

Have questions or a partnership idea? We'd love to hear from you.

Get in touch