Paper on Semantic Type Extraction for Data Lakes accepted to SIGMOD 2023

Congratulations to Sven Langenecker and Carsten Binnig!

2023/05/15

Sven Langenecker will present the paper “Steered Training Data Generation for Learned Semantic Type Detection” at SIGMOD 2023 in Seattle.

In this paper, we introduce STEER to adapt learned semantic type extraction approaches to a new, unseen data lake.

STEER provides a data programming framework for semantic labeling which is used to generate new labeled training data with minimal overhead. At its core,

STEER comes with a novel training data generation procedure called Steered-Labeling that can generate high quality training data not only for non-numeric but also for numerical columns. With this generated training data STEER is able to fine-tune existing learned semantic type extraction models. We evaluate our approach on four different data lakes and show that we can significantly improve the performance of two different types of learned models across all data lakes.