Skip to content
AI.info

Research

Kyrtos: A methodology for automatic deep analysis of graphic charts with curves in technical documents

Overview Research area: Computer vision and pattern recognition for document analysis — specifically the automatic recognition and deep understanding of graphic charts (curves) embedded in technical d

Kyrtos: A methodology for automatic deep analysis of graphic charts with curves in technical documents
arXiv
2602.09337
Published
2026-02-10
Authors
Michail S. Alexiou, Nikolaos G. Bourbakis

AI summary

Overview

Research area: Computer vision and pattern recognition for document analysis — specifically the automatic recognition and deep understanding of graphic charts (curves) embedded in technical documents, as part of the broader goal of Deep Understanding of Technical Documents (DUTD).

Technical level: Advanced. The paper combines unsupervised clustering (hierarchical clustering, K-Means, elbow method), formal language theory (grammar, alphabet, production rules), attributed graph representation, and Stochastic Petri-net (SPN) modeling.

Scope in one sentence: Kyrtos is a two-module methodology — a chart recognition module and a chart analysis module — that converts curves in chart images into straight line-segments, extracts their structural and behavioral relations using a purpose-built formal language, and expresses those relations as attributed graphs, natural language sentences, and SPN graphs.

What This Paper Is About

Technical documents contain charts whose knowledge is largely locked in image form, and existing chart-recognition methods tend to rely on restrictive heuristics (for example, assuming one-pixel-wide lines or only simple bar charts), or they assume the chart has already been converted into a data series by external tools. Kyrtos addresses this by automatically recognizing curves of arbitrary color, width, direction and continuity, and then analyzing them deeply enough to recover behavioral features such as direction, trend, growth and decay. The goal is to capture the visual knowledge illustrated in chart images, deduce internal structural and behavioral associations, and represent them in a common form — SPN graphs — that supports deeper document understanding.

Key Contributions

The paper states four main contributions:

  1. An automated method for recognizing and analyzing graphics in technical documents, overcoming structural restrictions and analyzing curves regardless of their characteristics (width, color, direction, expansion).
  2. Kyrtos integrates unsupervised learning with formal and statistical logic to improve analysis accuracy and provide deeper understanding of chart information.
  3. The introduction of the Kyrtos formal language, which bridges retrieved relations and maps them into SPNs — extending the earlier Glossa formal language.
  4. The conversion of chart associations into SPN format for enhanced understanding and recreation of chart functionality.

The paper also positions itself against prior work: attributed graphs have been used for engineering diagrams and tables before, but the authors state they have not been adapted for the representation of curves prior to this work.

Main Findings

  • Recognition without heuristic width/color rules: Kyrtos is described as invariant to heuristic rules about width, coloring, direction and expansion of chart curves, because it selects representative pixel-level features rather than assuming fixed curve properties.
  • Unevenness points as the structural unit: Curves are dissected at points where their direction changes. The authors build on an existing unevenness detection formulation and add clustering to handle curves thicker than the 1–2 pixel case that earlier methods (e.g., Nair et al.) were limited to.
  • Clustering thresholds: Hierarchical clustering uses upper and lower pixel-distance thresholds determined through trial and error — detailed clusters at 3 pixels (upper limit) and good clusters at 4 pixels (lower limit). These bound an iterative K-Means, whose cluster count is chosen by the elbow technique, with K-Means run five times per iteration and the result with the lowest distortion score retained.
  • Four chart types identified: (1) solid or partitioned curves in different colors, (2) partitioned curves of the same color with varying shapes, (3) bar graphics, and (4) pie charts. The paper focuses only on type 1 as a proof of concept for the methodology.
  • Attributed graph structure: Extracted line-segments become nodes with starting and ending pixel positions as attributes; connection, parallelism and intersection are arc labels. Connection arcs are labeled with angles in degrees, intersection arcs with 2D intersection locations, and parallelism arcs with the common slope value.
  • Behavioral classification: Growth analysis classifies curves as "exponential growth", "linear growth", "exponential decay", "linear decay", or "steady" (slope remaining zero across multiple segments, i.e., parallel to the x-axis).
  • Three merging rules: The first removes worst-case anomalies (endpoint y-axis distance of 1 pixel and x-axis distance of at most the curve's width minus 1 pixel); the second discards overlapping segments whose slope difference is within a threshold derived from that same worst-case ratio; the third merges two consecutive segments of the same curve that are both parallel to a common segment from a different curve.
  • Evaluation approach: The paper reports evaluating accuracy by measuring the structural similarity between input chart curves and the approximations generated by Kyrtos, for charts with multiple functions. The numeric evaluation results are presented in Section 7, which is not included in the available text, so no accuracy, similarity, or benchmark figures can be reported here.

Methodology in Plain English

The pipeline has two modules.

Recognition. Chart images are first separated from their axis information. Because standard segmentation and binarization cause miscoloring, missing edges or loss of entire curves, the authors use contrast and sharpen filters together with the HSV color space in what they call the HSV homogeneity filter, which replaces shades of a color with that color's dominant value. Pixels are grouped by dominant value to distinguish axis, grid and curves; the image is converted back to RGB. PyTesseract OCR reads axis values and labels, and the centers of axis elements' bounding boxes are used to associate recognized content with the original image.

Each curve's pixel trail is then parsed to find "unevenness points" — places where direction changes, detected using a slope equation and an unevenness criterion derived from the subimage's height and width. Because thick curves generate many candidate points for the same region, the points are clustered (hierarchical clustering to bound the search, then iterative K-Means with an elbow-based cluster count), and each cluster's middle-point — located mid-cluster with respect to x-axis positions, and explicitly distinguished from the centroid — becomes a segment endpoint. Pixels that share a curve's color but lie farther from it than a distance threshold equal to the curve's width are discarded as noise.

Analysis. Recognized segments are first condensed with the three merging rules. Intersections between segments of different curves are computed by solving a system of parametric line equations, accepting a solution only when both parameters fall between 0 and 1. Relations are then encoded in the Kyrtos formal language, whose grammar is defined as G = (VN, VT, GR, F) with non-terminals {Q, SL, C, GR} and terminals {V, @, !}, and whose operators include "><" for intersection, "=" for parallelism, "~" for connection, "@" for "and", and "!" for "or" (used when multiple interpretations of the same image are possible, such as partitioned curves of the same color).

These relations become attributed graphs. To link a curve's behavioral meaning to actual axis values, the method stores an 81x81 matrix around each segment midpoint (a size chosen by trial and error) and slides it across the original chart image to find matches — a simplification of a gradient-based approach by Lowe et al. Graph relations are then converted into natural language using Agent-Verb-Patient kernels, with verbs such as "is connected to", "is parallel to", "is intersecting with", and "illustrates exponential/linear growth/decay". Finally, these sentences map into SPN kernels — places as states, transitions as actions, firing rates controlling token flow over time. The resulting SPN graphs are used to regenerate the recognized curves, which are compared against the original chart image by structural similarity.

Why This Matters

Kyrtos targets the gap between shallow chart parsing and genuine comprehension of what a chart communicates. Most prior chart work either constrains the input (one-pixel lines, simple bars, pre-converted data series) or stops at detection rather than behavioral interpretation. By producing attributed graphs, natural language descriptions and SPN graphs from the same extracted structure, Kyrtos aims to make chart knowledge comparable and combinable with the other modalities of a technical document — tables, diagrams and text.

Real-world applications, based on the domains the paper situates itself in:

  • Automated document reverse engineering — extracting knowledge from the large volume of accumulated published research papers, which the paper identifies as the driving motivation.
  • Content-based document retrieval — using chart layout and structural characteristics, which the paper notes are converted into fuzzy attributed graphs for graph-matching retrieval.
  • Chart accessibility — the paper discusses related work on making charts accessible to visually impaired users through textual descriptions, a role the natural language output of Kyrtos directly serves.
  • Chart question answering and summarization — the paper cites transformer- and LLM-based chart comprehension and QA frameworks as adjacent to its goals.

Industry relevance: Any sector that produces or consumes technical documentation with plotted data — engineering, manufacturing (the paper cites control chart pattern recognition as a related problem), scientific publishing, and document-processing software vendors — stands to benefit from turning static chart images into structured, queryable, behavior-annotated representations. The SPN output is notable because it models functional behavior over time, not just geometry.

Future Directions

  • Extending beyond chart type 1. The paper explicitly frames its focus on solid or partitioned curves in different colors as a proof of concept, leaving partitioned curves of the same color, bar graphics and pie charts for separate, type-specific handling.
  • Detailed and broader evaluation. The reported approach measures structural similarity between input curves and Kyrtos approximations; Section 7's results are where quantitative evidence would live, and expanded benchmarking across chart types remains open.
  • Resolving ambiguous interpretations. The "!" operator exists because some images admit multiple valid readings; how reliably the system selects among them is an open question.
  • Holistic document understanding. The authors' stated broader goal is dissecting a technical document into all its modalities and associating them, so integrating Kyrtos with the table and text analysis pipelines (including the prior Glossa-based text-to-SPN work) is the natural continuation.

Target Audience

Researchers in document analysis, chart recognition and graphics understanding; pattern recognition and computer vision practitioners working on multimodal document intelligence; and anyone building systems for chart reverse engineering, chart summarization, or vision-language question answering over figures. Readers interested in formal-language and graph-based knowledge representation for visual data will find the Kyrtos grammar and its SPN mapping particularly relevant. A background in pattern recognition and formal languages is helpful, given the paper's level of technical detail.

Authors’ abstract

Deep Understanding of Technical Documents (DUTD) has become a very attractive field with great potential due to large amounts of accumulated documents and the valuable knowledge contained in them. In addition, the holistic understanding of technical documents depends on the accurate analysis of its particular modalities, such as graphics, tables, diagrams, text, etc. and their associations. In this paper, we introduce the Kyrtos methodology for the automatic recognition and analysis of charts with curves in graphics images of technical documents. The recognition processing part adopts a clustering based approach to recognize middle-points that delimit the line-segments that construct the illustrated curves. The analysis processing part parses the extracted line-segments of curves to capture behavioral features such as direction, trend and etc. These associations assist the conversion of recognized segments' relations into attributed graphs, for the preservation of the curves' structural characteristics. The graph relations are also are expressed into natural language (NL) text sentences, enriching the document's text and facilitating their conversion into Stochastic Petri-net (SPN) graphs, which depict the internal functionality represented in the chart image. Extensive evaluation results demonstrate the accuracy of Kyrtos' recognition and analysis methods by measuring the structural similarity between input chart curves and the approximations generated by Kyrtos for charts with multiple functions.

Read the original paper