Industrial-Instruction: An End-to-End Framework for Building Instruction-Tuning and Benchmark Datasets from Industrial Technical ReportsExplained for Beginners
Parsa Bakhtiari, Hassan Bashiri, Alireza Khalilipour +2 more
Abstract
Industrial technical reports contain high-value knowledge for maintenance, troubleshooting, and product engineering, but their heterogeneous structure (dense prose, specifications, tables) makes them difficult to index and reason over with standard retrieval and QA pipelines, and no public instruction-tuning or benchmark datasets are built from such documents. We address this gap with Industrial-Instruction, contributing (i) two open QA datasets built from real industrial technical reports and (ii) the end-to-end pipeline that produces them. Using 906 public Panasonic documents (7,525 pages), we apply layout-aware extraction, build a semantic retrieval index, and synthesize multiple-choice QA grounded in retrieved evidence under five query-document relationships (irrelevant retrieval, single-/multi-document support, single-/multi-document answer). After filtering an initial 23.9k generated samples, each dataset provides approximately 13.6k QA pairs with source documents and a held-out benchmark split. Fine-tuning small open LLMs (under 10B parameters) improves Set-Match Accuracy from 28.5% to 42.0% and F1 from 46.6% to 63.5% on the Panasonic benchmark. We release two parallel versions built by the same pipeline: one generated with the open-weight Qwen3-30B-A3B-Instruct model and one with the closed, API-based Claude-Opus-4.6 model, enabling a direct comparison of open- versus frontier-model data generation. The Claude-Opus-4.6 dataset yields a cleaner raw corpus and larger fine-tuning gains, at roughly two orders of magnitude higher cost. MMLU evaluation shows models trained on the Claude-Opus-4.6 data retain essentially all general knowledge, versus a small but measurable forgetting effect for the Qwen-generated data. Together, these datasets and pipeline offer a practical, reproducible path toward scalable industrial benchmarks and training data from real-world documentation.
Industrial-Instruction: Turning Technical Manuals into Teaching Data
The Problem
Imagine you are a maintenance engineer troubleshooting a failing industrial machine at 3 AM. You have a stack of PDF manuals on your desk, each hundreds of pages long. The machine’s error code points to a vague electrical issue, and you need to know which sensor to check first.
In the Industry 4.0 era, this scenario plays out millions of times a day. Industrial technical reports—maintenance manuals, specification sheets, failure analyses—are treasure troves of specialized knowledge. But they have a structural problem: they are messy. Dense prose, intricate tables, diagrams, and footnotes are all jumbled together in PDF format. Standard search engines and question-answering systems, which thrive on clean, linear text, struggle to make sense of them.
More importantly, until now, there were no public datasets built from these documents. AI researchers have plenty of benchmarks like SQuAD or MMLU, but nothing that captures the specific reality of industrial knowledge: the way tables encode specifications, the way a single failure mode might be described across multiple sections, or the necessity of ignoring irrelevant documents without hallucinating an answer.
This paper, Industrial-Instruction, tackles exactly this gap. The authors argue that to build useful AI for industry, we need instruction-tuning data that comes from actual industrial paperwork, not just internet text. They create two datasets from 906 publicly available Panasonic technical documents (7,525 pages total) and, crucially, they open-source the pipeline that creates them. This means other researchers can reproduce the work, adapting it to their own company’s manuals.
How It Works (The Technical Mechanics)
The paper describes a sophisticated end-to-end pipeline, but the core idea is accessible if we think of it as a "data assembly line."
1. Breaking the PDFs (Layout-Aware Extraction)
First, the raw PDFs need to be read. PDFs are notoriously difficult because they store text positions rather than a clean text stream. The authors use Dots.OCR, a vision-language model, to convert the PDF pages into Markdown—a format that preserves structure.
Think of this step like a highly advanced scanning process. Unlike simple OCR (Optical Character Recognition) which just reads text left-to-right, Dots.OCR identifies different "regions" on a page: this is a table, this is a heading, this is a footnote. It outputs the content in Markdown, which is friendly for large language models (LLMs).
2. Building the "Brain" (Semantic Indexing)
Once the text is extracted, the system needs to find information. The authors use a retrieval approach common in RAG (Retrieval-Augmented Generation) systems. They chunk the text into segments and convert each chunk into a "vector"—a numerical representation of meaning using the EmbeddingGemma model.
If you’ve ever used a search engine and it found a document even though you used different words than the author, that’s vector search. Here, the system can find a relevant paragraph even if the user’s query uses different terminology than the manual. They store these vectors in FAISS, a library designed for fast similarity search.
3. Generating the Q&A (The Synthetic Data Step)
This is the heart of the contribution. The pipeline automates the creation of question-answer pairs. It works like this:
-
The system retrieves a set of documents relevant to a topic.
-
It presents those documents to a powerful LLM (either the open-weight Qwen3-30B-A3B-Instruct or the closed Claude-Opus-4.6).
-
The LLM is prompted to generate a multiple-choice question based on those documents, but with a twist. The prompt explicitly tells the model to simulate five different "retrieval scenarios":
- Useless Document: The docs retrieved have nothing to do with the answer. The model must learn not to invent an answer.
- Single Document Support: One doc has clues, but the answer isn't written outright.
- Multi-Document Support: Clues are scattered across several docs; the model must connect them.
- Single-Document Answer: The answer is hiding in plain sight in one doc.
- Multi-Document Answer: The answer requires piecing together info from multiple docs.
-
Filtering: The initial run generates about 23,910 samples. A rule-based filter throws out samples where the multiple-choice options were poorly formed (e.g., just "A, B, C" without descriptions). This leaves roughly 13,600 high-quality QA pairs per dataset.
An analogy for software engineers: Imagine a version control system for knowledge. The pipeline takes the "source code" (the PDF manuals), runs a "build" step (the OCR and indexing), and then a "test generation" step (the LLM prompting). The output is a curated set of test cases that verify if a model can actually understand the code.
Key Results & Benchmarks
The real proof is in the fine-tuning. The authors take small open LLMs (models with under 10 billion parameters) and fine-tune them on their new datasets. They then test performance on the Panasonic benchmark (a subset of the data held out during creation).
Here are the headline numbers:
- Baseline: Without any fine-tuning, these small models achieve only 28.5% Set-Match Accuracy and 46.6% F1 on the industrial questions. This reflects the difficulty of the domain.
- After Fine-Tuning (Qwen-generated data): Accuracy jumps to 42.0% and F1 to 63.5%.
- After Fine-Tuning (Claude-Opus-4.6-generated data): The gains are even larger—Accuracy reaches 45.2% and F1 reaches 68.1%.
Translating the numbers: Fine-tuning on these industrial datasets improves a small model's ability to answer technical questions correctly by roughly 14 to 18 percentage points. In practical terms, if a model was getting about 1 in 3 questions right, it now gets about 2 in 5 right. That is a massive improvement for a small model running on limited compute.
Furthermore, the paper compares the two dataset versions:
- The Claude-Opus-4.6 dataset yields a "cleaner raw corpus" and larger fine-tuning gains.
- However, this comes at a cost: the paper notes the Claude data generation cost is "roughly two orders of magnitude higher" than the Qwen data. (This is a common trade-off: frontier closed models often produce higher quality synthetic data, but at a steep price).
There is also a fascinating MMLU (Massive Multitask Language Understanding) result. MMLU tests general knowledge (e.g., "What is the capital of X?" or "Solve this physics problem"). The authors fine-tuned models and then re-took the MMLU test.
- Models trained on the Claude-Opus-4.6 data retained essentially all their general knowledge.
- Models trained on the Qwen-generated data showed a "small but measurable forgetting effect," meaning they lost a tiny bit of general ability to specialize in industrial text.
Why It Matters (Key Takeaways)
- A Reproducible Pipeline for Any Company: The biggest takeaway is the pipeline itself. Most companies have reams of internal documentation. This paper provides the "recipe" to turn that documentation into training data. You don't need to hire an army of annotators; you can use this framework to generate QA pairs from your own manuals.
- Small Models Can Learn Industrial Skills: The results debunk the idea that only massive, billion-dollar models can handle industry. Fine-tuning models under 10B parameters on this data yields significant gains. This makes deployable AI feasible for industrial settings where compute resources are constrained.
- The Cost-Quality Trade-off is Real: The comparison between Qwen and Claude highlights a practical dilemma for AI teams. If you have a generous budget, the Claude-generated data gives better results. If you are cost-conscious, the Qwen-generated data still offers substantial improvement over a baseline, especially given the two-orders-of-magnitude cost difference.
- Robustness to Retrieval Noise: By explicitly training on "useless documents" and "multi-document" scenarios, the framework teaches models how to handle real-world retrieval failures. In a production RAG system, the retriever will sometimes pull the wrong doc; models trained on this data are less likely to confidently hallucinate an answer from noise.
Summary
Industrial-Instruction is a practical bridge between the world of academic AI benchmarks and the messy reality of industrial documentation. By releasing two high-quality datasets and the open-source pipeline that generated them, the authors give the research community a tool to build better, more specialized AI for maintenance, engineering, and troubleshooting. The key insight is that with the right extraction and generation pipeline, even small language models can develop "industrial common sense," closing the gap between general AI and specialized engineering expertise.
Want to understand AI papers like this from scratch?
Follow the free AI Learning Roadmap →