Extended Seminar - AI for Data Management

This seminar is about how AI can be used for data management. This year, we have reworked the seminar to focus on Semantic Data Systems that extend relational databases with AI operators to process multi-modal data. The course starts with a mini lecture series to provide the necessary background for the student presentations and practical tasks that follow.

  • Part 1 (Paper Presentations): You will present a publication from related work on semantic query processing, its optimizations, and benchmarks. Through your presentation and those of your fellow students, you will become experts on semantic data systems.
  • Part 2 (Implementation): You will implement a small semantic query processing engine and come up with your own novel ideas to optimize its accuracy and execution costs.

Organization

Last offered Winter Semester (26/27)
Lecturer Prof. Carsten Binnig
Assistants Jan-Micha Bodensohn, Liane Vogel, Zixuan Chen
Contact
Examination See Moodle
Kick-Off October 13th 2026, 9:50-11:30 AM (S306/146)

Course Infos

Below, you find some general information about the seminar. For all information regarding this year’s seminar (including important dates), please check the Moodle course linked above. Also make sure that you are registered in TUCaN.

Prerequisites:

Familiarity with databases, machine learning, and Large Language Models. Proficiency in Python.

Seminar Topic:

Despite decades of research and development, traditional relational database systems remain constricted to a handful of simple data types such as strings, dates, and numerical values. The queries they execute typically build on a small set of well-understood relational operators such as comparison filters, equi-joins, and numerical aggregations. By contrast, data today is often multi-modal, taking the form of images, videos, audio, and long-form documents. Queries have changed as well, from simple lookups and aggregations to complex analyses that often require multiple steps to derive new insights from raw data.

To keep up with these demands, a recent line of research focuses on so-called Semantic Data Systems that extend relational databases with AI operators to process multi-modal data. For example, a semantic join between a table of invoices and a table of pay statements may determine which payment has paid for which invoice, and a semantic aggregation can summarize salient arguments across a long list of legal documents.

While cloud database providers such as Snowflake, Databricks, and Google BigQuery already offer at least some basic AI operators such as semantic maps and filters, the development of semantic query processing engines has opened up a new frontier in database research. Since AI operators typically build on Large Language Models (LLMs), they significantly increase the cost of query processing, with some operators such as semantic joins remaining prohibitively expensive at scale. Moreover, their accuracy plays an important role since they are defined in natural language instead of a formal algebra.

This seminar is designed to introduce students to the foundational concepts of using AI for data management, with a special focus this year on Semantic Data Systems.