Trained from Scratch
Adriaticum was trained from random initialization rather than fine-tuned from an existing base model. As an approximately one-billion parameter transformer, its lineage progressed through scratch pretraining, regional language adaptation, and targeted instruction tuning without inheriting external model weights.
Linguistic Structure & Distinction
The model is shaped specifically for Serbian, Croatian, and Bosnian. While these languages share substantial grammatical structure and mutual intelligibility, they preserve distinct standard vocabularies, orthographies, and regional usage. Adriaticum respects these distinct linguistic identities without conflating them into a single generic approximation.
Dual-Script Support
Adriaticum provides coverage across both Latin and Cyrillic scripts. The system handles standard orthographies, technical vocabulary, and regional expressions directly across both alphabets without artificial transliteration steps.
Documented Training Provenance
Pretraining and language adaptation drew from curated regional corpora—including parliamentary proceedings such as ParlaMint and regional web text from CLASSLA-web—alongside multilingual datasets. Detailed licensing notices and source attributions are published in our Third-Party Notices.
Access & Availability
The conversational workspace is the primary interface for interacting with Adriaticum. Public API execution is not currently available.