NetBench: domain adaptation for networking LLMs
What does domain adaptation teach a language model about high-performance networking?
- Status
- Released v1.1.0 · paper in preparation
- Authors
- Licence
- Code Apache-2.0 · data CC-BY-4.0
The question
Teams adapt general language models to specialist fields by fine-tuning them, continuing their pre-training on domain text, or giving them retrieval over a document store. NetBench asks what each of those methods actually changes about what a model knows of high-performance networking: bulk data transfer, congestion control, bottleneck diagnosis, and the parameters that govern them.
The benchmark
HPN-QA has 242 open-ended questions, 233 of them scored, generated from 175 open-access publications. Each answer is graded against a reference on correctness, completeness, clarity, and conciseness. Every item carries its source evidence with DOI and licence, and a three-model calibration ladder assigns its difficulty.
The study
Eight open-weight models from 1B to 27B parameters answer the same items under up to eight adaptation variants: the instruct baseline, LoRA and full supervised fine-tuning, continued pre-training, and retrieval. Three frontier API models answer them too, as baselines. Two independent LLM judges score every answer, and a two-expert human study validates the primary judge. The design and statistical protocol were fixed before any results were analysed.
- Collect papersAcquire open-access PDFs and remove duplicates.
- Build the corpusConvert PDFs to cleaned text.
- Generate dataProduce the HPN-QA benchmark and instruction-tuning data.
- Adapt and answerFine-tune, pre-train, or add retrieval, then generate answers.
- Judge and analyseTwo LLM judges score; analysis code builds every table and figure.
Reproducibility
make reproduce regenerates every CSV, table, and figure behind the paper from the committed judged outputs, and fails if anything comes back different. It runs on every push. Source PDFs, the text corpus, and model weights are not redistributed; the repository records their checksums and explains which steps can be re-run and which cannot.
Status
Version 1.1.0 was released on 30 September 2026. The paper describing the study is in preparation, and its results will be linked here once it is public.
Cite as: N. H. Neom, R. M. Swargo, M. Arifuzzaman. NetBench, v1.1.0. Zenodo, 2026. doi:10.5281/zenodo.21892017
