Nieb Hasan Neom

PhD student in Computer Science
Missouri S&T

Seeking Summer 2027 research and software engineering internships.

← Research

NetBench: domain adaptation for networking LLMs

What does domain adaptation teach a language model about high-performance networking?

Status
Released v1.1.0 · paper in preparation
Authors

N. H. Neom, R. M. Swargo, M. Arifuzzaman

Licence
Code Apache-2.0 · data CC-BY-4.0

The question

Teams adapt general language models to specialist fields by fine-tuning them, continuing their pre-training on domain text, or giving them retrieval over a document store. NetBench asks what each of those methods actually changes about what a model knows of high-performance networking: bulk data transfer, congestion control, bottleneck diagnosis, and the parameters that govern them.

The benchmark

HPN-QA has 242 open-ended questions, 233 of them scored, generated from 175 open-access publications. Each answer is graded against a reference on correctness, completeness, clarity, and conciseness. Every item carries its source evidence with DOI and licence, and a three-model calibration ladder assigns its difficulty.

The study

Eight open-weight models from 1B to 27B parameters answer the same items under up to eight adaptation variants: the instruct baseline, LoRA and full supervised fine-tuning, continued pre-training, and retrieval. Three frontier API models answer them too, as baselines. Two independent LLM judges score every answer, and a two-expert human study validates the primary judge. The design and statistical protocol were fixed before any results were analysed.

  1. Collect papersAcquire open-access PDFs and remove duplicates.
  2. Build the corpusConvert PDFs to cleaned text.
  3. Generate dataProduce the HPN-QA benchmark and instruction-tuning data.
  4. Adapt and answerFine-tune, pre-train, or add retrieval, then generate answers.
  5. Judge and analyseTwo LLM judges score; analysis code builds every table and figure.
One data pipeline feeds six self-contained modules. The direct and retrieval variants share judge prompts, rubric, and output schema, so their results compare directly.

Reproducibility

make reproduce regenerates every CSV, table, and figure behind the paper from the committed judged outputs, and fails if anything comes back different. It runs on every push. Source PDFs, the text corpus, and model weights are not redistributed; the repository records their checksums and explains which steps can be re-run and which cannot.

Status

Version 1.1.0 was released on 30 September 2026. The paper describing the study is in preparation, and its results will be linked here once it is public.

Cite as: N. H. Neom, R. M. Swargo, M. Arifuzzaman. NetBench, v1.1.0. Zenodo, 2026. doi:10.5281/zenodo.21892017