{"id":6,"date":"2023-10-18T15:10:19","date_gmt":"2023-10-18T13:10:19","guid":{"rendered":"https:\/\/projets.litislab.fr\/finlam\/?page_id=6"},"modified":"2025-06-13T17:00:33","modified_gmt":"2025-06-13T15:00:33","slug":"finlam","status":"publish","type":"page","link":"https:\/\/projets.litislab.fr\/finlam\/","title":{"rendered":"FINLAM : Foundation INtegrated models for Libraries Archives and Museum"},"content":{"rendered":"\n<div class=\"wp-block-columns is-layout-flex wp-container-core-columns-is-layout-3a88641f wp-block-columns-is-layout-flex\">\n<div class=\"wp-block-column is-layout-flow wp-block-column-is-layout-flow\">\n<p class=\"wp-block-paragraph\">The digital transformation of libraries, which has been based on OCR (Optical Character Recognition) technology for more than 20 years, faces certain limitations both in terms of quality, due to the diversity of the collections and the limitations of OCR technology, and in terms of added value, due to a lack of structuring and high-level indexing. Named entity extraction is still little used because it mobilises language processing technologies, which were not very adaptable until recently. More generally, the semantic indexing of collections is underdeveloped and integrated with metadata. We propose to develop multimodal models (text + image) for the extraction of information from collections of digitised documents in large libraries. The literature shows that work in this direction is still underdeveloped, and that it is mainly aimed at processing commercial documents (invoices etc\u2026).<\/p>\n<\/div>\n\n\n\n<div class=\"wp-block-column is-layout-flow wp-block-column-is-layout-flow\">\n<figure class=\"wp-block-image aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"190\" height=\"249\" src=\"https:\/\/projets.litislab.fr\/finlam\/wp-content\/uploads\/sites\/10\/2025\/04\/logo-FINLAM.png\" alt=\"\" class=\"wp-image-32\" \/><\/figure>\n<\/div>\n\n\n\n<div class=\"wp-block-column is-layout-flow wp-block-column-is-layout-flow\">\n<p class=\"wp-block-paragraph\"><strong><a href=\"https:\/\/anr.fr\/Projet-ANR-23-IAS1-0007\">ANR programme FINLAM<\/a> <\/strong>relies on the expertise of <strong><a href=\"https:\/\/www.litislab.fr\">LITIS<\/a><\/strong> to study the most relevant multimodal architectures to integrate the language knowledge conveyed by the large language models developed recently and to study the modalities of specialisation\/adaptation of these models in conjunction with the learning of a generic optical encoder, benefiting from the annotated collections available at the <strong><a href=\"https:\/\/www.bnf.fr\/fr\">BnF<\/a><\/strong>. User interaction will be considered according to different scenarios of closed and open queries. <strong><a href=\"https:\/\/teklia.com\/fr\/\">TEKLIA<\/a><\/strong> will mobilise its expertise to prepare the data and carry out model integration experiments and the deployment of its production chain on some targeted corpora. Specific user interaction scenarios will be specified and implemented, and will give rise to original experiments conducted in collaboration with the BnF operators and users. The BnF will be responsible for evaluating the performance of the proposed solutions quantitatively and qualitatively in terms of ergonomics, usability, and acceptability by the users.<\/p>\n<\/div>\n<\/div>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"114\" src=\"https:\/\/projets.litislab.fr\/finlam\/wp-content\/uploads\/sites\/10\/2025\/05\/Pied-de-page-1024x114.png\" alt=\"\" class=\"wp-image-72\" srcset=\"https:\/\/projets.litislab.fr\/finlam\/wp-content\/uploads\/sites\/10\/2025\/05\/Pied-de-page-1024x114.png 1024w, https:\/\/projets.litislab.fr\/finlam\/wp-content\/uploads\/sites\/10\/2025\/05\/Pied-de-page-300x33.png 300w, https:\/\/projets.litislab.fr\/finlam\/wp-content\/uploads\/sites\/10\/2025\/05\/Pied-de-page-768x85.png 768w, https:\/\/projets.litislab.fr\/finlam\/wp-content\/uploads\/sites\/10\/2025\/05\/Pied-de-page-1536x170.png 1536w, https:\/\/projets.litislab.fr\/finlam\/wp-content\/uploads\/sites\/10\/2025\/05\/Pied-de-page-2048x227.png 2048w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n","protected":false},"excerpt":{"rendered":"<p>The digital transformation of libraries, which has been based on OCR (Optical Character Recognition) technology for more than 20 years, faces certain limitations both in terms of quality, due to the diversity of the collections and the limitations of OCR technology, and in terms of added value, due to a lack of structuring and high-level [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":0,"parent":0,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"footnotes":""},"class_list":["post-6","page","type-page","status-publish","hentry"],"_links":{"self":[{"href":"https:\/\/projets.litislab.fr\/finlam\/wp-json\/wp\/v2\/pages\/6","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/projets.litislab.fr\/finlam\/wp-json\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/projets.litislab.fr\/finlam\/wp-json\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/projets.litislab.fr\/finlam\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/projets.litislab.fr\/finlam\/wp-json\/wp\/v2\/comments?post=6"}],"version-history":[{"count":10,"href":"https:\/\/projets.litislab.fr\/finlam\/wp-json\/wp\/v2\/pages\/6\/revisions"}],"predecessor-version":[{"id":108,"href":"https:\/\/projets.litislab.fr\/finlam\/wp-json\/wp\/v2\/pages\/6\/revisions\/108"}],"wp:attachment":[{"href":"https:\/\/projets.litislab.fr\/finlam\/wp-json\/wp\/v2\/media?parent=6"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}