Adriaticum

A language model built for the linguistic context of the Balkans

Adriaticum is Mijatovic's core language model, trained from scratch with particular attention to Serbian, Croatian, and Bosnian language data and their shared and distinct linguistic structures.

01

Trained from Scratch

Adriaticum was trained from random initialization rather than fine-tuned from an existing base model. As an approximately one-billion parameter transformer, its lineage progressed through scratch pretraining, regional language adaptation, and targeted instruction tuning without inheriting external model weights.

02

Linguistic Structure & Distinction

The model is shaped specifically for Serbian, Croatian, and Bosnian. While these languages share substantial grammatical structure and mutual intelligibility, they preserve distinct standard vocabularies, orthographies, and regional usage. Adriaticum respects these distinct linguistic identities without conflating them into a single generic approximation.

03

Dual-Script Support

Adriaticum provides coverage across both Latin and Cyrillic scripts. The system handles standard orthographies, technical vocabulary, and regional expressions directly across both alphabets without artificial transliteration steps.

04

Documented Training Provenance

Pretraining and language adaptation drew from curated regional corpora—including parliamentary proceedings such as ParlaMint and regional web text from CLASSLA-web—alongside multilingual datasets. Detailed licensing notices and source attributions are published in our Third-Party Notices.

Access & Availability

The conversational workspace is the primary interface for interacting with Adriaticum. Public API execution is not currently available.