Veyra

Veyra

Startup Launched Aug 2026
Share:
Veyra social preview
Preview of Veyra

The Story

We built Veyra to create the first Indonesian language model that truly understands Indonesian structure. Traditional models break down words with affixes, losing meaning. Our NMU framework keeps Indonesian words intact—one character, one ID—so the model learns from the language's authentic roots, not translations.

AI Overview

AI-generated

Indonesian language models have historically struggled with a fundamental structural challenge: the language's rich system of affixes and word modifications gets fractured by tokenization approaches designed for English. Veyra, a 75-million-parameter language model, addresses this problem through an alternative architecture that treats Indonesian on its own terms.

The core innovation is the NMU framework—ninmeni meaning unit—which assigns a fixed ID to each character. Rather than breaking down inflected words into subword tokens, the model ingests complete words with all their affixes intact. This preserves the semantic structure that Indonesian speakers naturally recognize, training the model on authentic linguistic patterns instead of reconstructed approximations. The approach reflects a deliberate design philosophy: that language models should learn from a language's native roots, not translated or adapted paradigms.

The model was built entirely from scratch using this framework, with the 75M parameter size chosen as a conscious validation point. Developers can verify model capabilities through systematic evaluation rather than selective demonstration. The focus remains deliberately narrow: Indonesian language performance shapes corpus selection, evaluation priorities, and development direction.

Veyra shows unexpected capability beyond its primary scope. Despite no explicit English training, it generates grammatically sound English sentences—an artifact of how the NMU character-level encoding treats Latin characters universally across both languages. Developers treat this as an observable phenomenon rather than a marketed feature, documenting it while continuing to investigate the mechanism.

The product positions itself as a tool for builders working specifically with Indonesian language applications. By rejecting the usual transfer-learning approach that adapts English-trained models for other languages, it offers a model trained in the way Indonesian actually works. This appeals to developers seeking more authentic language understanding for Indonesian, teams building primarily for Indonesian-speaking users, and researchers interested in non-English-centric language model design.

The emphasis on documented evaluation and transparency about capabilities—including what remains unverified—indicates a research-first mindset that prioritizes credibility over marketing claims. No pricing information appears in available materials, suggesting this may be an open-source or research-stage project focused on validating the NMU framework before commercial deployment. That restraint itself signals maturity: the commitment to prove the approach works before scaling operations.

Key Features

Character-Level Tokenization

Uses the NMU framework to assign fixed IDs to each character, preserving complete words with all affixes intact instead of fragmenting them.

Indonesian Language Optimization

Built from scratch specifically for Indonesian with corpus selection and evaluation priorities centered on authentic linguistic patterns.

75M Parameter Validation

Deliberately sized to enable systematic evaluation of model capabilities rather than selective demonstration.

Morphological Preservation

Maintains semantic structure that Indonesian speakers naturally recognize by treating inflected words as complete units.

Unexpected English Capability

Generates grammatically sound English sentences despite no explicit English training, due to how character-level encoding treats Latin characters universally.

Use Cases

  1. 1

    Indonesian Application Developers

    Build language applications optimized for Indonesian linguistic structure rather than adapted from English-trained models.

  2. 2

    Indonesian-Speaking Markets

    Create products for Indonesian-speaking users with authentic language understanding native to Indonesian.

  3. 3

    Language Model Researchers

    Study non-English-centric language model design and alternative tokenization approaches for morphologically rich languages.

FAQ

How does Veyra handle Indonesian word affixes?
Veyra treats complete words with all affixes intact through its NMU framework, rather than fragmenting them through English-designed tokenization.
Can Veyra generate English text?
Yes, despite no explicit English training, Veyra generates grammatically sound English as an artifact of how its character-level encoding treats Latin characters universally.
What is the model parameter size?
Veyra is a 75-million-parameter model, deliberately sized as a validation point to enable systematic evaluation of capabilities.
Is Veyra open source?
No pricing information appears in available materials, suggesting it may be open-source or a research-stage project focused on framework validation.

Tech Stack & Tags

Discussion

No comments yet — be the first!

Join the conversation — sign up to comment.

Sign up free
0

Community Support

Boost this project on Sell With boost

Meet the Founder

Launch your own

Getting discovered has never been this beautiful.

Submit a Startup