IEEE VIS 2025 Content: Large Language Models for Transforming Categorical Data to Interpretable Feature Vectors

Large Language Models for Transforming Categorical Data to Interpretable Feature Vectors

Karim Huesmann -

Lars Linsen -

Image not found
Screen-reader Accessible PDF

Room: Hall E2

Keywords

Vectors, Data visualization, Frequency measurement, Encoding, Image color analysis, Automobiles, Large language models, Data models, Semantics, Computational modeling

Abstract

When analyzing heterogeneous data comprising numerical and categorical attributes, it is common to treat the different data types separately or transform the categorical attributes to numerical ones. The transformation has the advantage of facilitating an integrated multi-variate analysis of all attributes. We propose a novel technique for transforming categorical data into interpretable numerical feature vectors using Large Language Models (LLMs). The LLMs are used to identify the categorical attributes’ main characteristics and assign numerical values to these characteristics, thus generating a multi-dimensional feature vector. The transformation can be computed fully automatically, but due to the interpretability of the characteristics, it can also be adjusted intuitively by an end user. We provide a respective interactive tool that aims to validate and possibly improve the AI-generated outputs. Having transformed a categorical attribute, we propose novel methods for ordering and color-coding the categories based on the similarities of the feature vectors.