This site requires Cookies enabled in your browser for login.
Updating ...
WaterNet Home
WaterNet
for
pour le
Canada
Menu
WaterNet
Home
GWFO
Home
Catalogue
Master Index
Data
Centre
Collections
X
Defaults
Select All
Websites
X
Global Water Futures Observatories (GWFO) Global Water Futures (GWF) Global Institute for Water Security (GIWS) International Network of Alpine Research Catchment Hydrology
Legacy Research Programs
X
Changing Cold Regions Network (CCRN) Drought Research Initiative (DRI) International Network of Alpine Research Catchment Hydrology (Legacy Site) Improving Processes & Parameterization for Prediction in Cold Regions Hydrology (IP3) The Mackenzie Global Energy and Water Cycle Experiment (GEWEX) Study (MAGS)
Legacy sites
Map
Utilities
X
Account Settings Create a New Record Record List Alias List Editor
Edit Data Centre
Data Types
. . .
X
Clear
Select All
Advanced Search
Go to Top⇡
Related items loading ...
Fetching Chart ...
Publication Additional Information Download
Publication Type
Thesis
Authorship
Oladipo, Akintunde
Title
Scaling pre-training data & language models for african languages
Year
2024
Publication Outlet
UWSpace - Theses
DOI
https://hdl.handle.net/10012/20872
Citation
Oladipo, Akintunde (2024) Scaling pre-training data & language models for african languages, UWSpace - Theses, https://hdl.handle.net/10012/20872
Abstract
Recent advancements in language models, particularly for high-resource languages, have not been paralleled in low-resource languages spoken across Africa. This thesis addresses this gap by scaling pre-training data and developing improved language models for African languages. We introduce Wura, a high-quality, document-level pre-training dataset encompassing 16 African languages along with four high-resource languages commonly spoken on the continent: Arabic, English, French, and Portuguese. Leveraging Wura, we pre-train new versions of the AfriBERTa (encoder-only) and AfriTeVa (encoder-decoder) model families. These new models demonstrate superior performance across a variety of natural language understanding and generation tasks compared to existing baselines. Notably, AfriTeVa V2 Large (1B) stands as the largest sequence-to-sequence model pre-trained for African languages to date. Our methodology includes a meticulous three-stage curation process for Wura --- auditing and filtering existing web crawls, initiating new web crawls, and integrating existing language resources. The experimental setup and evaluation encompass tasks like text classification, information retrieval, translation, summarization, and cross-lingual question answering. Our new models outperform their predecessors and other established models, even those with significantly more parameters, highlighting the efficacy of high-quality pre-training data. Furthermore, we study the generalization of our models to languages not deliberately included in their pre-training data.
Program Affiliations
GWF: Global Water Futures
Publication Stage
Published
Download Links
https://uwspace.uwaterloo.ca/bitstreams/0eb3cf2b-5660-492e-8be6-6f4d4e82f5e1/download
© 2026 - WaterNet Version 2026-07-16
Global Water Futures Observatories
Powered by
G W F Net
T-2024-12-19-V18Mg640CIUe7XXxcp2t6bw Publication 1.0