Vision-Language Models for Network Model-Single-Line Diagram Consistency Verification: A Feasibility Study
DOI:
https://doi.org/10.31224/8231Abstract
Single-line diagrams (SLDs) are central to power system planning and operations and must remain consistent with the machine-readable network models used by downstream applications -- a manual, hard-to-scale reconciliation made increasingly pressing as network models are revised ever more frequently. Vision-Language Models (VLMs) could automate this, but their application to SLDs remains under-explored, in part because power systems is a low-resource domain lacking expert-validated benchmarks. We conduct a feasibility study on VLMs for SLD Consistency Verification: identifying discrepancies between an SLD and its text-based network model representations. We first find that general-purpose VLMs display underwhelming performances. Hypothesizing that this stems from an inability to attend to the relevant component within a dense diagram, we evaluate whether contrastive VLMs (e.g., CLIP) can localize queried components; finding that they too struggle, whereas traditional Optical Character Recognition (OCR) proves surprisingly effective, we propose an OCR-anchored cropping strategy. Integrating this with VLMs substantially improves performance, though still below expectations in safety-critical scenarios. As a preliminary feasibility study, we hope our findings direct attention toward resources spanning two future pathways for GenAI in power systems: (1) curating expert-validated benchmark datasets and (2) developing a power-systems foundation VLM.
Downloads
Downloads
Posted
License
Copyright (c) 2026 Charles Alba, Dean Miller, Rishabh Jain, Karthik Kumar, Seong Lok Choi, Fei Ding, Benjamin Kroposki

This work is licensed under a Creative Commons Attribution 4.0 International License.