ProBackend
ai medical diagnostics remote healthcare
1 hour ago5 min read

Adapting TeenyTinyLlama for Brazilian Portuguese Medical QA: The Doctor Llama Project

Research on Doctor Llama, a project featuring compact models fine-tuned from TeenyTinyLlama on Brazilian Portuguese medical Q&A data.

Training small language models on specialized regional data is notoriously tricky, especially when you step outside English. Most open-source architectures are saturated with Anglo-centric datasets, leaving low-resource languages like Brazilian Portuguese stranded with mediocre zero-shot translations or massive models that are impossible to run locally without specialized hardware infrastructure. That is precisely why Mariana Moreira dos Santos undertook the Doctor Llama project as her biomedical informatics course completion thesis at the Federal University of Paraná (UFPR). By taking the compact TeenyTinyLlama architecture and fine-tuning it exclusively on Portuguese medical questions, this work tackles both linguistic scarcity and domain specificity head-on.

The Compact Language Model Challenge in Brazilian Portuguese

Medical natural language processing in Portuguese has historically lagged behind English due to a scarcity of open, curated datasets and the heavy computational costs associated with large-scale pre-training. Hospitals and research groups in Brazil often rely on proprietary models or direct translations of English benchmarks, which frequently fail to capture local clinical terminology, pharmacological nomenclature, and regional healthcare nuances.

Smaller models—such as the sub-half-billion parameter tiers—offer an appealing alternative for edge computing and low-resource academic settings, and the cost case for running such small language models locally has only strengthened as enterprise teams look to cut cloud inference spending. However, raw base models trained on general web crawls possess zero innate clinical reasoning. Without targeted fine-tuning, they flounder when asked about anatomical structures, diagnostic protocols, or symptomatology in Portuguese. The Doctor Llama initiative directly addresses this gap by exploring how far supervised fine-tuning can push miniature architectures when trained strictly on curated medical question-and-answer pairs.

Anatomy of the Doctor Llama Project

The initiative centers around two distinct model scales derived from the TeenyTinyLlama family: a 160-million parameter variant and a 460-million parameter variant, alongside chat-tuned iterations (Doctor Llama Collection). Rather than throwing billions of parameters at the problem, dos Santos wanted to understand the granular mechanics of adapting miniature open weights to clinical terminology in Brazilian Portuguese without introducing catastrophic forgetting. That question mirrors a broader reassessment underway across medical AI, where evaluations such as comparisons of Llama models of different sizes on medical tasks suggest parameter count alone is a poor predictor of clinical performance.

The underlying datasets published within the Hugging Face collection form the backbone of this experiment. The primary training corpus, medicine-training-pt, consists entirely of medical questions formulated in Brazilian Portuguese. This strict filtering forces the network to grapple with local clinical phrasing, pharmacological terms, and diagnostic queries without diluting its focus on general web noise. Paired with medicine-evaluation-pt, this dataset setup provides a clean, rigorous benchmark to measure how well miniature causal language models internalize specialized jargon during supervised training cycles.

Training Infrastructure and Hyperparameters

Getting miniature architectures to converge on specialized text without overfitting or breaking syntax requires careful optimization. The fine-tuning runs leveraged an NVIDIA A100-SXM4-40GB GPU, completing the training cycle in approximately five hours per configuration.

Under the hood, the training pipeline employed the torch.optim.AdamW optimizer configured with a learning rate of 1e-5, an epsilon of 1e-8, and 1,000 warmup steps. Models were trained across 4 epochs with a batch size of 8, operating within a 2,000-token context window (Doctor Llama 460m). This setup hits a deliberate sweet spot: aggressive enough to imprint complex medical vocabulary onto lightweight weights, yet controlled enough to prevent the model from degrading its underlying grammar during backpropagation.

Evaluation Benchmarks and Perplexity Gains

Empirical numbers tell the real story of parameter adaptation here. Across the board, fine-tuning on medicine-training-pt yielded sharp reductions in both perplexity and evaluation loss when compared against the untuned base TeenyTinyLlama models (Doctor Llama 160m).

For the 160m tier, the base TeenyTinyLlama model logged a perplexity of 22.51 and an evaluation loss of 3.11. Doctor Llama 160m dropped those figures to 15.68 and 2.75, respectively. The larger 460m architecture saw similarly impressive gains, starting from a baseline perplexity of 13.09 and an evaluation loss of 2.57, and improving down to 10.94 and 2.39. Even the chat variants showed marked progress: the base TeenyTinyLlama 460m Chat stood at 21.22 perplexity with a 3.05 loss, while Doctor Llama Chat brought that down to 11.13 perplexity and a 2.41 loss.

These empirical metrics prove conclusively that even sub-half-billion parameter models can capture complex domain distributions when fed clean, targeted regional training sets rather than undifferentiated web scraps. They also highlight an open gap: unlike the large general-purpose systems examined on open community leaderboards for medical LLMs, Doctor Llama's gains are measured primarily in perplexity and loss, not in standardized clinical multiple-choice benchmarks.

Limitations, Safety, and Future Research

We need to be entirely transparent about what Doctor Llama is not. It is an academic research project examining fine-tuning dynamics in Brazilian Portuguese, not a commercial diagnostic tool ready for clinical deployment in hospitals or clinics.

The model card explicitly warns against human-facing interactions or production use without extensive local risk and bias assessments. Like many compact language models, Doctor Llama inherits inherent tendencies toward hallucinations, historical social biases, and occasional unreliable statements. Furthermore, its proficiency is strictly bound to standard Brazilian Portuguese; feeding it prompts in English or other languages leads to immediate comprehension failures.

Yet, as an open-source contribution distributed under the Apache 2.0 license, Doctor Llama offers a remarkably valuable blueprint. It demonstrates that researchers don't need massive enterprise cluster budgets to build domain-specific language tools in under-resourced languages—they need disciplined dataset curation, clear experimental design, and open science sharing.

the compact language model challenge in brazilian portuguese

More blogs