Engineering Drawing OCR

Turn drawings into structured BOM data

Read title blocks, part tables and revision histories straight off CAD sheets and engineering PDFs — including scanned and legacy archives. Runs air-gapped on your own hardware, and every value traces back to the spot on the drawing it came from.

100%
On-premise & air-gapped
CPU
No GPU cluster required

The Problem

Every part your team sources begins life on a drawing. But the drawing is a picture, and your PLM needs records. So somebody reads the title block, retypes the part numbers, transcribes the component table, and checks the revision — by hand, per sheet, across thousands of sheets.

Generic OCR does not solve this. It reads text in straight lines and hands back a wall of characters, with no idea that a value belongs to the Material field, that a number is rotated ninety degrees along a dimension, or that a row belongs to the parts table rather than the notes. And it cannot be used on the drawings that matter most, because CAD sheets are the last thing a manufacturer will send to a public cloud API.

Deployed at
Maruti Suzuki

Automated parts list ingestion from engineering drawings, running inside their own network — enabling faster engineering review cycles.

What It Extracts

Structured records, not a block of text.

How It Works

Three steps. Every value verifiable against the original sheet.

1

Ingest

Drop in CAD exports, PDFs or scanned sheets — single drawings or an entire legacy archive. Mixed quality and mixed formats are expected, not a problem.

2

Read the layout

Custom spatial vision layers locate the title block, the parts table and the revision rows, read rotated and vertically aligned text, and rebuild the BOM hierarchy rather than returning loose characters.

3

Match & export

Parts are matched against your PLM or ERP catalogue and exported as structured records. Anything uncertain is flagged for a human, with a one-click jump to the exact region of the drawing.

Why This System

Air-gapped by design

Runs entirely behind your firewall with no external API calls. Your drawings — the most sensitive IP you hold — never touch the public internet.

Traceable, not trusted

Every extracted value links to a bounding box on the original sheet. A reviewer verifies by looking at the drawing, not by taking the model's word for it.

Runs on ordinary hardware

Mixed-precision INT4/INT8 quantisation and weight pruning let the vision and layout models run on standard office CPUs — no GPU cluster, viable at remote plants.

Built for real drawings

Rotated text, dense parts tables, stamps, signatures, skewed scans and decades-old archive sheets. Trained on how drawings actually look, not idealised samples.

Generic OCR vs. Engineering Drawing OCR

AspectGeneric OCRGrokking
OutputUnstructured textStructured BOM records
Rotated textFrequently missedRead natively
Table structureLostPreserved with hierarchy
DeploymentUsually cloud APIAir-gapped on-premise
VerificationNo audit trailBounding box to source
HardwareGPU or cloudStandard office CPU

FAQ

What can the system extract from an engineering drawing?

Title block fields such as drawing number, title, scale, material and sheet number; the itemised component or parts table; revision history rows; and part numbers wherever they appear on the sheet. These are assembled into a structured Bill of Materials rather than returned as loose text.

Does it work on scanned drawings and legacy archives?

Yes. The pipeline uses custom-trained vision transformers and grid layout detectors rather than plain line-based OCR, so it handles scanned sheets, low-resolution images, skewed pages, stamps and handwritten annotations. Converting legacy drawing archives into structured data is one of the most common uses.

How is this different from standard OCR?

Standard OCR reads text in horizontal lines and returns a block of characters. Engineering drawings need spatial understanding: text rotated ninety degrees along a dimension, values that belong to a specific cell of a parts table, and a title block whose meaning depends on position. The system reads layout, geometry and relationships, then outputs structured records instead of raw text.

Can it push data into our PLM or ERP?

Yes. Extracted BOM trees are reconstructed in a hierarchy compatible with PLM and ERP databases, and parts are matched against your existing catalogue so the output lands as usable records rather than a spreadsheet someone has to re-key.

What hardware does it need?

Standard office workstations. We apply mixed-precision INT4 and INT8 quantisation and weight pruning to the underlying vision and layout models so they run on ordinary CPUs, which means no dedicated GPU cluster and no cloud dependency. This is what makes deployment viable at remote plants and branch facilities.

Do our drawings ever leave our network?

No. The system is designed for fully air-gapped deployment. It runs behind your firewall with no external API calls, which matters because CAD sheets are among the most sensitive intellectual property a manufacturer holds.

Related

Other systems in the Grokking suite.

RFP Evaluation AI Sales & Quotation Engine

Book a Demo

Send us a sample drawing and we will show you exactly what comes back. Tell us about your archive and our team will get in touch shortly.

Chat on WhatsApp