l for for Lana thank you very much for being here uh I'm I'm huge fan of yours of course and uh I'm so um looking forward for uh this uh interview I understand what you say good good it's good things I'm I'm saying I'm talking about Good Things um so uh for and thank you luchana again for AC accepting the invitation and participating on this Congress and sharing your your knowledge thank you so much in Vancouver Canada Washington for for thank you um so let's start with uh the ongoing read research uh at interpares AI
uh could you share with us some of the research that interpares AI is currently undertaking and what have been some key findings or challenges that the project has encounter thus far okay so first of all I would like to thank you uh for the presentation and for the kind words that you have said and I would like to say hello to to all those who will listen to what I have to say and apologize if I am not brilliant at times it's it's much easier to follow PowerPoint than to improvise answers to questions but anyway
um so uh the interp Paris sayi is the fifth phase of inter Paris and uh its purpose really is to try and understand how artificial intelligence tools can be used in archival functions to carry out arival functions uh but to do so uh by uh involving arist in developing uh artificial intelligence tools so the purpose is to have archival Concepts and archival methodologies embedded in whatever archival tool is generated um so of course uh the biggest challenge that we have encountered before talking about any finding has been to talk to each other because the people
involved in this research my co director himself is the Canada chair for artificial intelligence is's a computer scientist he doesn't understand archives uh and so so do alpha of the researchers in the entire team we have about 200 researchers from about 30 not about precisely 37 countries uh and uh they are all part of other universities or archival institutions uh but generally speaking half of the researchers are not archist of Records managers and that means that we had to first of all try to speak the same language so the most important general study we have
done has been a terminology database when a computer scientist talk about data it refers to any digital content so when mam says I need more data he's not talking about data we're talking about in in archival science when we say data we refer to the smallest indivisible piece of information within a record to computer scientist Records books Christmas cards anything you mention is Digital Data as long as it is digital okay so we have developed tutorials for archist which are actually posted on the website uh accessible to everybody which are sort of Elementary explanation of
how AI Works uh how the neural networks work etc uh also arches taught uh diplomatics and archival science at the beginning of the various plenaries to um computer scientists digital forensic people lawyers we have got many disciplines involved because of course we need ethics Specialists we need uh law Specialists Etc so anyway uh that has been the biggest challenge it took us about a couple of years to be able to talk to each other sure that the others know what we're talking about uh it's not an easy thing so uh that has been the challenge
is mostly overcome um at least uh there is no longer the certain on the part of computer scientists that they know what we're talking about and at least they ask the questions and as we do as we archist records managers do of them the findings well the most important finding is we're not going to have a y archives anytime soon it's too complex we are working on it we have developed tools I will talk about them later but uh as the situation is now there is no way we can entrust our function to uh artificial
intelligence tools uh they can help and whatever they do can be verified and I can show it as I talk about the various studies thank you for sharing that luchana now we're going to jump to we know interp parties is heavily has been relying over and over in several case studies and I I believe this is part of the richness of your findings and and how you present you know um the results of your research so could you provide examples of case studies or General Studies that is working on so we have a total of
about 50 studies wow uh the General Studies are the studies that apply to everything so not a specific environment or not a specific um field or function uh so yes we did tons of annotated literature that we have on the website the terminology database is key to what we do uh all the terms used are defined first in the way we use it within the research project then they are defined the way they are used in the various disciplines so that people can make the connection and also how they were used in the previous interpares
projects because they remain relevant to this one um we have the tutorials that I have mentioned uh we have study that is um almost complete but of course will not be till the end no study will be completed till the end of the research which is at the end of 26 uh but we have enough that we can actually say this is how it is so um AI Literacy for records managers at archist uh there are two big studies on this one is by Moses rocken in with Brazilian archist but he works in Portugal uh
and the other one is from uh Richard darias Hernandez uh who is a colleague of ours in the school of information uh we have a meta study believe it or not so the meta study which is directed by Ken tibo is actually an analysis of all the studies together categorizing them in order to identify the the hols where it is that we're not going and where it is that we keep going not getting anywhere and what are the methodologies that are most useful for what studies so for example um we have uh natural language understanding
uh that he is a technology that he is important to understand speech when we have oral records or when we have audiovisual records and to also understand uh text uh for various purposes uh we have natural language um generation uh tools which are for uh summarization to summarize to make a summary of whatever thing we need to summarize um also for translation uh and we we need transcriptions and we need translations and I will get there uh we have natural language processing uh for classification uh and there is there is a model success the the
success on classification which is one of the studies carried out ched by the Malaysian group uh um one uh you know it is it has a moderate success but it entirely depends on how good the classification system is to start with because if you have a strong classification system and you have plenty of content records which have been classified according to that system then you can easily develop a model that reflects both the system and the way it is used so you have a high probability of accuracy you don't have it if you don't have
a classification system and that's what it is uh you know AI has been used just records thrown to it and say organize it wait a second you can't organize records if you don't know the entire context and the entire context is totally embedded in a classification system so you need to have a classification system possibly linked to a attention and disposition system which would allow also for extending this specific tool to selection right and to an appraisal function okay so um we also uh use uh can use it for arrangement and description so there is
a study on this one that has been carried out now for two years by a very large Italian team working with the Norwegian Technology Group uh which has always claimed that they can do it so they have worked together to develop uh a proper way of recognizing the original arrangement of the records and end of describing and at the end of the project the the Norwegian team has said it's too complicated we're not going to do this okay that's why I said is not very close to F AI because it is not that they can't
it costs money to create a system that has that level of accur y that is needed so that's why we need to do things a little bit at a time from the creation time because if we have the things that are properly classified organized Etc then all the rest comes along but not when we have the traditional way of dealing with the records okay so um and then of course we have been very successful in developing uh an AI tool for detoxification it's called Grill Lama detoxification means that they will find in appropriate language the
the the tool will find in appropriate language and substitute it with proper language um so for example um no example comes to mind in this moment but might later um so this this is the kind of things that we are working in in two primary Labs uh the key lab is the one of UBC it is in the Linguistics department is ched by Muhammad who is my director Muhammad Abdul maid um and it has has several um students from various backgrounds who are working on the development of these tools the other uh lab is at
the University of machata and it's mostly a vision uh uh um Vision Computing um lab and I will describe one study by conducted by the in this lab which resulted in the development of a tool when we talk about tools um so um as it regards uh the studies that are case studies and are more advanced uh we can look at the UNESCO study so the UNESCO study is carried out by the people in UNESCO as well as the lab of mamed and Linguistics um and so UNESCO has digitized um not long ago um I
think here 16,000 tapes of interviews uh made between 1950 and 1980 uh with all the big names of the time in the cultural sector in the political sector in the educational sector Etc and yes they're digitized they have got no metadata nobody knows what is in them other than those are all the important persons were asked important questions and answer in 75 different languages identified and often uh well one quarter of the tapes um have multiple languages in them so we want to create we want an AI tool that creates metadata by identifying the persons
uh the subject matter the time uh and all the other factors that are relevant for each tape so uh we are developing this tool for is basically natural language processing but it is much more difficult when it is an interview than when it is a text that is written because the text that is written the paragraphs help you uh the punctuation helps you uh in the in a in a speech transcription you have access you don't have just different languages you have accents you have different way of uh of um behaving in your answers uh
hesitations uh repetition you have look at me the way I talk try to transcribe what I say right so it is not simple but it is advancing there is quite a bit of progress in this um and um the the major difficulty is that AI is good with the main uh Western languages it is not good with any African language and for this reason mamed has developed an entire set of tools that actually is able to deal with African languages so that's an achievement uh and it is not good with minor languages and definitely not
with dialects so what we have here is we start with the transcription of the text uh automatic transcription of the text of the speech right um we have the set of 50 metadata per record that UNESCO requires and we have this connection so that what is found in the speech can be linked to so it recognize names it recognize stes it recognize certain things but the most important thing is when you using diplomatics which is recognized by the eye people working on this as the best tool possible why no matter how you interview somebody in
the radio or in a private interview or in a conference how do you start what's the event what are the names what the subject matter what are the dates what's the time right so that's the structure of the record when you look at all those records they all have the same structure and the difference is when there are radio programs in which case all that information is at the end instead of being at the beginning it doesn't matter it's still all together so diplomatics has shown to be useful not all in this study if you
take the study that we have on the private information so the actual uh um recognition uh in uh records of uh private personal information uh it is it is useful and it is being used in a study to um to actually label the records you know that in order order to be able to use content we need to be able to label it so to put a little square around it and say what that is okay well that's what diplomatics has done from the beginning that's the order that's the address that's the the uh the
writer that's the date that's the place that's the action that's the subject matter Etc so the ability to label automatically all the material so then the relevant information can be ex structed whatever that is in this case of this study is the Privacy related personal privacy related information but it is also used to diplomatics to identify the types of records that are likely to contain personal information and the types that are not likely to according to their diplomatic form so these two you see in every team there are as many artificial intelligence people as there
are archist so we are we are carrying out a very important educational work in teaching the AI expert how to understand records how to understand actions what act is this so therefore what does it involve okay now when we talk about privacy we are not just studying uh in the traditional trying to develop something right we're also looking at what exists and how it can be used because AI is not a new thing it has existed for decades so shouldn't we just take whatever exist and works and so we we in fact there is a
study that goes through the chain of preservation actions of the previous interpares where you have in this chain all the actions involved along the life of the record right so what this study does is to identify existing AI for each one of those actions that we carried out because why should we waste our time if they work so what we do when it comes to access there are the Privacy announcement tools so one of these tools is called distant reading what it does is you have no access to anything the this tool has access to
all the material and is's able to show it to you without showing you uh the actual information so it only gives you what you need and nothing else okay then uh however leaking of information using this tool is not that difficult so that's why it is by itself is not such a great idea it does the job of if somebody really want to get some private information still can uh what we do is and this is Vick Lemo that is guiding this this research is to combine distant reading with price privacy preserving Federated machine learning
that I yet to read it because it never comes all together okay privacy preserving Federated machine learning and what it means is that you have a bunch of organizations each one one uses this tool without making available the material to the others is only the come that is common to everybody so that protects privacy much better um other studies that we have um I talked about rangement description access um images Okay so this that is Guided by Jessica bushie and uh it is about yeah uh images uh how they are used uh AI IM AI
generated images uh she has investigated how with our team how they're used in medicine how they're used by the police uh how they're used the education environment Etc and see what the consequences are um one study that is complementary to this really is the study on veracity so you know that in archival science we never cared about what is in the record is true or false doesn't matter it's still the record right it's up to the researcher to figure out if the story told in the record is true or false with artificial intelligence it is
a big of a problem to to delegate this to the user because while we can say in the public sector uh we can control all of this and make sure that we create records that are uh reliable in terms of content uh not just in terms of authorship and process of creation um also the public bodies receive records from private individuals and how do we verify that actually this are not forgeries so this is a study that just started I mean six months ago so at this point we are all at the end of the
literature review we're not very Advanced so well these are a few examples there are many more studies there are studies on there have been lots of questionnaires that have gone out and we have results from those um related to um have gone to users uh to ask users about what they use if anything that is artificial intelligence related or what or how they would like to access the material so we we have several studies uh related also to automatic digitization um Etc but for now this will give you an idea uh if you want the
complete list you go on the website in Paris trust ai.org and you click on studies and there is The List complete of every study that is being carried out in some cases if you go to dissemination on that website you actually find reports on the studies on the status uh of the various studies uh for example we have at least three refereed articles of bad data so there are plenty of Publications that have been coming out thank you luchana so at the end of your uh of your answer you mentioned paradata so could you elaborate
on how paradata creation and management can be effectively integrated into archival practices to ensure the long-term accountability of AI systems okay so um first of all often accountability is confused with explainability and those are two very different things because explainability just tells you what the algorithm is and what it does uh what is the expected outcome uh it doesn't assign responsibility and the fact is that you may say well if something goes wrong the responsibility is of Google open AI whatever no they will tell you as they have in many uh lawsuits that have occurred
uh especially in the UK they will tell you you B the two you WR read the instructions you use it we have nothing to do with it so what you want uh archist and Records managers to do if they use AI is to cover their back and that's what paradata is for okay so paradata is not metadata metadata is about the records it tells you what the record is and now you retrieve it Etc par data is about activities so it's about uh who is carrying out the activity of using AI how this person has
been trained uh what the policy of the institution is why has chosen that specific uh AI Tool uh what do you want to achieve with this tool what the procedure has been every step of using it uh what the logs are what outcomes is so basically um there are two types of paradata uh technical paradata which are well the model as it has been developed uh the evaluation of the model and its performance so all the tests done all the metrics the logs that are generated in the process um the the data set for training
as separate from the data set that is actually used to get results um the training parameters and if you buy the tool rather than developing the tool you want to have all the vendor documentation all the versioning information Etc if you're dealing with uh or okay that's that's the technical paradata as it regards organizational paradata uh you need the AI policy of the institution the design plans of the institution the employee training all the ethical considerations that have been um expressly identified um the the process of implementation and all the regulatory requirements that are relevant
to that institution so under What legislation under what uh regulations the thing is um now the the general par ofata are about what is documented uh it's it's uh is about okay determining not only what the system is but what parts of the system you need to preserve because it's not enough that you say what you have used you have to keep that so you have to keep the code the model the algorithm the logic the executable the training data the test data the results of the test the validation data um and those are in
addition to Old organizational data and the procedural data right and then you need to have the operational data so let me explain this because this is important there are three types of AI system one type is one that doesn't involve human beings at all so basically it has sensors that collect the information the data transmitted to the control the control examines that and decides on the action and transmits this to the actuators they realize the action carrying it out and then the outcome goes back to the sensors so it is a loop that keeps going
and going okay in this Loop humans are not involved an example is uh the robot okay uh that does things from beginning to end the cleaner that can do it from beginning to end okay so that is no problem because that is all totally described as it is the second type of system is a recommendation system a recommendation system is one to which humans May react or not so basically does all the same thing as the other only the human receives the recommendation and may act upon it or not well think every time you click
on the internet on a product beauty product the next day you get 25 recommendations who do you think did that there is the artificial intelligence system that knows oh she is interested in this Tech so let's go and tror all the possibilities now the human is involved but marginally is not the serious thing the serious thing happens when actually the human and the system act at the same time so that is when in fact you get what we call shared agency which takes me to question four that's a name so um as you mentioned there
are specific issues that arise with shared agency a AI systems especially in the context of reservation would you elaborate on some of these issues and how they impact archival practices yes okay so let me give you examples of shared agency uh the airport in Vancouver we have uh passengers uh controls and guidance okay this is a joint action so there is a room somewhere with all screens that sees where the passengers are and where they should go through security or to a specific gate or whatever so controls the tra traffic of the passengers in the
airport however the job is actually done by an AI tool that tells the the passengers where to go depending on what flight they have to take or depending on whether security is too busy and they should go to some other part Etc so they act at the same time in the sense that the humans control what AI is doing and have the ability to correct or to act at the same time uh and to intervene in some way okay um and so it's it's complicated but not that complicated you get really complicated when you go
to do you remember the blueprints you know what the blueprints are okay the the actual drawings of buildings and all the details those are used by Engineers to build something right they are added to and after the building is complete they must be preserved for as long as the building exists and sometimes longer after okay so those are key records in a CD okay we don't have blueprints anymore what we have is what we call digital twins digital twins are actually digital images of the buildings at each moment as they change if I turn on
the eating system you can see warming up if I add something to a balcony I can see it changing so what's the problem with this one it's a dynamic environment there are time constraints for intervening you have to intervene within a specific time um there are different levels of autonomy uh the control uh it may change because when you have a shared agency when there is concurrent action of AI and human beings it may be the AI that corrects the human beings not only the human being that correct AI or they can act in conflict
or one can override the action of the other I mean think of the car right the car that are controlled by AI there is a point where you can say okay you drive and you don't do anything else right but then if it looks bad that there is an accident to happen it's it's who responsibility is to avoid it the AI or the individual who is driving it's the same thing with all these other with the digital twins so you have concurrent action that sometimes is overlapping sometimes is in Conflict sometimes is to correct sometimes
is to enance and it is IR real time so sensing perception acting it all happens in real time who responsibility is so you say well what do I care I am an archist I'm not going to use that system in an archives no but you have to acquire the records and that's the question you don't have blueprints anymore so what records do you acquire how do you intervene in those system to be able to freeze them at any given time so that you have the evidence of what has been going on and what's the code
of a certain uh consequence right so what we need to do here we go back to parad data okay the the paradata so you need not only the images at any given time you need the data at every predefined trigger point and so you you need and you need to follow this feedback loop model as you present the data at every trigger point so the temporal Dimension is important uh how often those things are uh those data are collected so you have to you have to maintain the relationship between all the processes as they appen
it's not simple now we have lots of media science people working on this in addition because we got we got by ourselves to deal with this but this is a big thing because it is not just for buildings is now basically for everything is done whatever is is created is developed Etc it is all done this way and we no longer have the drawings that we could go to so this becomes a serious preservation system um and so you know AI enabled system must be documented minute by minute or faster and of course uh that
is the primary basis to determine whether there is authenticity or not in the records that we we have because you may believe them looking at them but if you don't have the documentation that says what it is what who has intervened when and what the consequences of the interventions are been then you don't really have anything and when we talk about buildings it is a very serious thing I mean the entire constructions that happen at culton University is where the chair of this group works it's all like that it's it's all like that and you
feel the precarity of all this how precarious all all of this is so that's the difficult part uh so paradata is our concern in terms of if we use AI we have got to document what we do but they are also concerned in terms of when we preserve something like digital twins what data we need to be with them not just meta data but data about how the AI is functioning because the metadata is what the building is who is constructing it what the time is and all that sort of stuff but the par data
is how this is documented in the AI system and how the AI is showing things and whether it is late with respect to the action for example I mean whether the action is corrected by humans at any given time and these are the parad data these are not data about the record these are data about what has been happening in the use of AI from beginning to end so that's the big difference between the too okay I think I spoke too much no it's never enough right taana well you know so the last question um
what types of tools are being developed by interpares AI that can be practically applied in archival institutions and could you describe some of the the successes in this area you mentioned this in the beginning of this talk right so yeah yeah so I mentioned some of the tools that are developed in the natural language lab uh and I mean that the one about uh biases uh is you know one cute thing about this tool about the biases I have to say it's not very politically correct to say it but I have it this about political
correctness right it is a tool to identify bias so when the person uh who was the primary person responsible for developing the tool showed it to us all the researchers and showed us how the bias were reduced and how the the the balance was generated with good language Etc the part with bias was colored in pink and the part with the correct thing was colored in blue and I raised my hand and I said I think that there is a big bias in this [Laughter] image but that you know was not intended of course but
it seemed so funny to look at that thing anyway that is one of the tools now one tool that actually is complete and has been used uh was developed at the University of machera for the study on parchment records um so the purpose of the study was to develop a tool to identify the primary attributes of thousands of parchments uh that had been digitized but for which they had to create metadata and then you nothing okay so um they use computer vision uh which is a field of AI uh that enable systems to Der to
derive information from digital images and they yet to choose what the main identifier of this parchments could be and decided that would be the Signum now the Signum the Signum is a sign a fixed to patment by noties uh to identify them by hand uh it identifies them both at the beginning of the parchment and at the end where the signature is and what they did uh by um by doing this by taking images of the syum in every patment they actually were able to establish registry of notaries uh for all the medieval period uh
to attribute parchments um so even to organize them uh so it was extremely useful but the thing is that this um tool that is called peronet you can find it on the internet easily if you just type Panet um is useful for many other things because it is useful for um recognizing the system of writing of people of individual authors it recognizes and writing but also the the type of writing uh in so that can attribute authors to I mean writings to others also is good for analyzing The archival annotation behind the document you see
lots of paper documents on the back have annotations among which often is the original order it's it's the classification or the order in which they were kept and by doing that by taking pictures off the back of Records they actually are able to use this system to find the original order or to organize according to original classification Etc uh it also recognizes recurring images in large series of documents in which people may not know what kind of records are there and you see in those who know diplomatics the special signs when you have the letter
head often there is the name of an organization and then there is a symbol that represents the organization either before or after or on the side and this kind of uh image will allow actually to find out to the Creator if not the others of the records are uh when you have large amounts of Records you see lots of people are very concerned about AI is going to steal our job we is going to do the things that we want to do we don't want to do that we don't want to go through every single
record and find out who the other is so what a what we try to develop AI for is to make our work Pleasant so we do the fun part and we don't do the boring part like looking for that private information inside the letter that needs to be blanked out uh we don't want to do that so this is useful and it also recognize common patterns in drawings also for attributing drawings so computer vision is really a very important thing there is a lab in matchata that is precisely for that so we use these two
main Labs one in UBC and one in machata for different purposes uh because the one UBC is for languages the one in matchata is for images mostly so that is an instrument that has been extremely useful and uh we are in the process I think that the the tool that we're going to use for UNESCO is close to being ready um so here we are great thank you Lana um I don't know we don't have more questions but can I add one you can have a question I can't guarantee an answer no yeah it's now
it's more a matter of your opinion because you've you've been you know teaching for so many years and sometimes I notice that our you know the Professionals in our field sometimes are very resistant to Technologies especially the emergent ones um and I believe for example based on everything you said today uh we need to specialize ourselves like archivist they have to think you know like U becoming more and more digital because I don't think the threat is the emerging technology but the threat is the lack of knowledge to deal with them and you know take
Pro uh proper care of the records so I just wanted to hear your thought about this I find this question interesting because clearly you come from a different environment because the environment in North America is exactly the opposite they only want to learn the technology they definitely don't want to learn archival theory or archival diplomatics and I have to say that the people in this research project who have been enthusiastic about diplomatics are the computer scientists they are not archist they didn't even want to use it it's no because okay we know that it is
not helping they don't they don't start with the idea that if the theory is good it will work with everything okay it's just that doesn't so anyway yes we have those two studies about AI literacy that is for archist and Records managers it's not for others M and and and uh they they are pretty Advanced I mean there is I believe that in the dissemination uh part of the website there are some of the speeches presented about this right so you should you should look for that um yes absolutely it is necessary to know specifically
what you don't know uh that is when it is that you have to go to The Specialist and when it is that you have to control the process because you can't say oh I learn everything so I do it by myself no you don't all the work that related to digital must be done in collaboration by specialists in different areas so you don't but uh you must know where your knowledge is essential and where you have to let people make the choices you just and you know it's not just the tendency of the archist uh
not to do that because I have observed the computer scientist to hear one day about records and I think they understand them all okay so it is just um when I say then I see Cal by then and I said and I say perhaps you want me to read what you wrote before you send it out you know it some things these things must be done in collaboration you can't have an archist who write an entire article about AI in archives you can't have an AI person who does it yeah because it the depth of
knowledge necessary in each of the field is such that you can do it on your own it it doesn't say anything about your intelligence or about your willingness to learn it's just recognition of the depth of fields and I want our field to be respected as much as I want to respect the field of yeah AI experts yeah and that's really important great thank you Lana so tach do you have any questions I don't know if our time is out yeah well um I have plenty of questions here but um I can ask later because
uh actually um yeah the time is out um I want to thank Lana so much for your answers and for being here and for your kindness and thank you thank you so much well thank you very much for having me H I hope that I have not just confused you no not you're you know how to explain the things you know the examples are great so thank you so much [Music] [Applause] yeah all the best everyone and thank you thank you bye bye by hiow ciao