EHRSQL: A Practical Text-to-SQL Benchmark for Electronic Health Records

Cited 0 time in webofscience Cited 0 time in scopus
  • Hit : 51
  • Download : 0
DC FieldValueLanguage
dc.contributor.authorLee, Gyubokko
dc.contributor.authorHwang, Hyeonjiko
dc.contributor.authorBae, Seongsuko
dc.contributor.authorKwon, Yeonsuko
dc.contributor.authorShin, Woncheolko
dc.contributor.authorYang, Seongjunko
dc.contributor.authorSeo, Minjoonko
dc.contributor.authorKim, Jongyeupko
dc.contributor.authorChoi, Edwardko
dc.date.accessioned2023-09-13T08:03:26Z-
dc.date.available2023-09-13T08:03:26Z-
dc.date.created2023-09-13-
dc.date.created2023-09-13-
dc.date.created2023-09-13-
dc.date.created2023-09-13-
dc.date.issued2022-12-
dc.identifier.citationNeurIPS 2022-
dc.identifier.urihttp://hdl.handle.net/10203/312591-
dc.description.abstractWe present a new text-to-SQL dataset for electronic health records (EHRs). The utterances were collected from 222 hospital staff, including physicians, nurses, insurance review and health records teams, and more. To construct the QA dataset on structured EHR data, we conducted a poll at a university hospital and templatized the responses to create seed questions. Then, we manually linked them to two open-source EHR databases-MIMIC-III and eICU-and included them with various time expressions and held-out unanswerable questions in the dataset, which were all collected from the poll. Our dataset poses a unique set of challenges: the model needs to 1) generate SQL queries that reflect a wide range of needs in the hospital, including simple retrieval and complex operations such as calculating survival rate, 2) understand various time expressions to answer time-sensitive questions in healthcare, and 3) distinguish whether a given question is answerable or unanswerable based on the prediction confidence. We believe our dataset, EHRSQL, could serve as a practical benchmark to develop and assess QA models on structured EHR data and take one step further towards bridging the gap between text-to-SQL research and its real-life deployment in healthcare. EHRSQL is available at https://github.com/glee4810/EHRSQL.-
dc.languageEnglish-
dc.publisherNeural Information Processing Systems Foundation-
dc.titleEHRSQL: A Practical Text-to-SQL Benchmark for Electronic Health Records-
dc.typeConference-
dc.identifier.scopusid2-s2.0-85161171551-
dc.type.rimsCONF-
dc.citation.publicationnameNeurIPS 2022-
dc.identifier.conferencecountryUS-
dc.identifier.conferencelocationNew Orleans-
dc.contributor.localauthorSeo, Minjoon-
dc.contributor.nonIdAuthorHwang, Hyeonji-
dc.contributor.nonIdAuthorKwon, Yeonsu-
dc.contributor.nonIdAuthorShin, Woncheol-
dc.contributor.nonIdAuthorYang, Seongjun-
dc.contributor.nonIdAuthorKim, Jongyeup-
dc.contributor.nonIdAuthorChoi, Edward-
Appears in Collection
AI-Conference Papers(학술대회논문)
Files in This Item
There are no files associated with this item.

qr_code

  • mendeley

    citeulike


rss_1.0 rss_2.0 atom_1.0