We do a lot of document processing at OBLSK. Recently, we had the opportunity to dig in on parsing data from architectual drawings, more specifically, home floor plans. As with many other documentation challenges, we thought feeding the challenge through an AI pipeline to solve the problem would be the solution. We were wrong. This post will share an example of challenges you might face trying to use a vision llm to parse from a floor plan and show the path we took to get the data we needed.
For our example, let’s answer one simple question. . .How big is each room in the floorplan?
Every number in this post comes from one file, simple_floorplan.dxf. It is a 14-room single-story house, and each room is stored as a closed polyline. One method reads those polylines. The other three read a PNG render of the same file. Nothing else varies.
To begin, let’s provide some structure to the objects we will be dealing with. Each object has a name, a polygon as a list of [x, y] vertices, and an area computed from those vertices with the shoelace formula.
Joining a predicted polygon to a real room needs a rule, and the rule changes the answer. I used centroid containment. A prediction joins a room when the prediction’s centroid falls inside that room’s native polygon. Intersection-over-union would have scored one enormous blob as a partial hit on many rooms. Centroid containment scores it as exactly one wrong room, which is the honest reading.
The scoreboard, before the story
Four readings of one file, scored against one 14-room schedule with one join rule.
| Method | Input | Found |
|---|---|---|
| Native DXF polylines | CAD entities | 14 of 14 |
| Gemini 3.8 Flash | PNG render | 12 of 14 |
| CubiCasa5K | PNG render | 0 of 14 |
Finding rooms with Gemini
I started with a general-purpose multimodal LLM, Gemini 3.8 Flash.
Gemini found all 14 rooms. It placed every centroid inside the correct room.
There were some issues though: The Living Room is 612.3 square feet and Gemini called it 611.3, off by 0.2 percent. The Entry is 56.2 and came back 57.2, off by 1.8 percent. The Pantry is 28.9 square feet and came back 30.5, off by 5.5 percent. The Utility Room is 35.0 and came back 34.0, off by 2.9 percent

Finding rooms with CubiCasa
Next, I tested CubiCasa5K. It is a neural network trained specifically on architectural floor plans, making it the most promising candidate on paper.
It returned a single region labeled “Other rooms” covering most of the drawing. That blob measures 1,514.6 square feet, and the room its centroid lands in is a 168.2 square foot bedroom. It is nine times too large. One room joined, zero areas within tolerance, and thirteen rooms never located at all.

There is no partial credit to award.
Training distribution explains the failure. CubiCasa learned from marketing-style floor plans with shaded rooms, furniture symbols, and filled wall cavities. A bare CAD line drawing with no fill is a completely different visual domain. A floor-plan network is not a DXF parser, and assuming it can read CAD wireframes because its title contains “floor plan” is a mistake.
Native CAD geometry
The native pass is short enough to describe completely.
simple_floorplan.dxf holds 14 LWPOLYLINE entities, one closed polyline per room. A companion CSV gives each room a name and a coordinate point. We apply the shoelace formula to each polyline for its exact area, then join each CSV point to the polyline that contains it.
Native geometry recovered 14 of 14 rooms, all 14 names, and all 14 areas with zero measurement error. That is the schedule every other method was scored against: a 612.3 square foot Living Room, a 432.2 square foot Garage, three bedrooms at 191.3, 168.2, and 148.5, down to a 26.9 square foot walk-in closet, totalling 1,909 square feet. It is also the green outline in all three figures above.
simple_floorplan.dxf, its room mapping CSV, and the construction sheet that supplied the scale all come from CCWI/building-plan-viewer under MIT. That sheet originates in jscad/sample-files, also MIT. Having the room polygons published alongside the drawing enabled this comparison without manual labeling.
What this does not prove
One drawing is one drawing. simple_floorplan.dxf is small, clean, single-story, and its rooms are already closed polylines. That is the friendliest possible input for every method here, which is what makes the spread interesting. How these four methods behave on a real construction sheet is a separate measurement, and it is not in this post.
The square feet deserve their own caveat. The scale is derived, not measured off a physical print. It rests on a companion sheet’s dimension style declaring architectural units and on two of its overall dimensions matching the room-polygon span once exterior walls are subtracted. That is strong agreement, and it is still an inference from the file rather than a surveyed building. Nothing in the comparison depends on it, because relative error cancels the unit.
It is also not a claim that vision-language models lack utility on drawings. This render carries no text at all, so nothing here tests label reading. Gemini 3.8 Flash located all fourteen rooms and got 12 areas inside 2 percent on the friendliest drawing in the set. That is real. It is still not a 14-row schedule with zero measurement error, and it is one model on one PNG.
If you are evaluating a vendor on document work, the more general version of this is on our evaluation guide, which asks for denominated failures on hard document types rather than a single accuracy percentage.
The habit worth copying
The verdict matters less than the method that produced it.
Build the native schedule first and treat it as truth. Then read the same file every other way you are considering, and score each one with the same join rule and the same tolerance. A miss shows up as a specific row with a name and a number attached, rather than a vague impression of degraded quality.
The design conclusion follows directly from the measurements. When a pipeline must produce numbers derived from geometry, read the geometry. A PNG or a PDF is a picture of the geometry, and flattening to pixels discards the exact vectors you need, but sometimes the challenge is that you may only get your hands on the PNG or a PDF.
My thanks to Isaac Flath, whose writing on parsing text, tables, and structure out of PDFs prompted me to measure the adjacent case where the payload is closed geometry rather than prose.