Eigen RadarAI
Analysis

Small language models propose table features while experts check their values

Emrah Inan and Damla Oguz developed a workflow that uses small language models to propose additional features for tabular classification. Retrieved documents support the suggestions, while a domain expert filters candidates and checks external values. Tests on country and animal data evaluated this division of work: the language model proposed feature names rather than inventing the numbers used to train the classifier.

Artificial Intelligence··Night
An expert works standing at a laptop on a raised desk.

Models propose features, experts collect values

Emrah Inan and Damla Oguz of İzmir Institute of Technology, a technology university, developed a semi-automated workflow for adding features to classification tables. A small language model proposes feature names and explanations. Candidate filtering falls to a domain expert, who discards unsuitable or duplicate suggestions and checks values against external sources. The language model does not estimate those values; conventional machine-learning models train on the expanded table after collection.[1]

Search connects descriptions with documents

The pipeline extracts metadata from labels and schema descriptions, chooses some existing features and prepares queries for document searches. A sentence-representation model connects queries to a vector database, which stores numerical text representations. Web searches retrieve documents for candidate features. Experiments included Gemma3:4b and LLaMA3.2:3b. Commercial ChatGPT and Claude comparisons were excluded on grounds of the expense of large-dataset use.[1]

Country and animal tables test the workflow

Global Country Information supplied 195 samples and 34 features, while Animal Information supplied 205 samples and 15 features. The tasks classified life expectancy into four classes and animal social structure into 10. Their highest F1 scores were 0.383 and 0.848. The approximate processing times per query were 2.34 seconds and 3.21 seconds. Feature reviewers reached an agreement measure of 0.85. F1 measures the balance between precision and recall.[1]

References

  1. News sourceScientific Reportsİzmir team uses small language models to propose tabular features↩1↩2↩3