Testing AI enhancement possibilities (theme Ethnic Cleansing)

In line with CISI’s commitment toward exploring responsible usage of artificial intelligence in historical research, an obvious starting point is the book of its head. Written by your guinea pig, Vladimir Petrovic, the book Etničko čišćenje: geneza koncepta (Ethnic Cleansing: Origins of the Concept) was published in 2019. It had a limited outreach for several reasons. Firstly, it was written in Serbian language. Secondly, only 300 copies were printed, which is a standard in the Serbian academic publishing. Thirdly, the book underwent soft censorship, as then minister of culture banned its distribution to public libraries. Back in the days, I did what i could to counter those obstacles. In agreement with the publisher, I made intro and summary in English available on Academia.edu, where it scored miserable 227 views over five years, despite my sizable following. A paradox ensued — book’s reception was positive, especially in the international academia, as I summarized main findings in an article in English and repeatedly held lectures about this phenomenon at prestigious academic venues. The chapter in English made it to Wikipedia,but was also barely cited. So is the book, according to google scholar. Therefore, I have to admit my defeat – for over ten years I was intensely trying to find out how did the phrase ethnic cleansing emerge. I successfully documented that it did not appear in the context of the 1990s Yugoslav wars nor Kosovo riots in early 1980s, but in 1941, appearing almost simultaneously in Romania and on the ruins of wartime Yugoslavia. Yet, I was able to mediate my findings only to my closest colleagues, friends and family.

This disappointing result is not only my failure. It is indicative for limited outreach of scholarly production in the global semi-periphery. We have to face the fact that we barely partaking in global academic exchange. We are mostly writing for ourselves, citing each other out of curtesy often without reading. Enormous effort is being repeatedly wasted, serving only our vanity and academic advancement within an autarkic system we created. I have to pause and remember an alarming sentence of my first mentor, Professor Andrej Mitrovic, who in early 2000s opened a state-of-the-art conference in Belgrade with a disturbing thought: “Serbian historiography does not exist on a global map.” Since then, things were going from bad to worse. Despite efforts of individuals and institutions, due to various factors (limited funding and publish-or-perish challenge in Serbia, general disinterest in the humanities and loss of interest in the Balkans in the postwar period abroad) we are virtually impaired and disadvantaged. Even when we have something to say, we have no audience. Our books are on the shelves, mainly collecting dust. Back in 2020, when Institute for Contemporary History created its Digital Center, our ambition was to overturn that trend by a digital jumpstart. Did we achieve anything? I do not know, because we were too busy digitizing, and did not really set any measurable yardstick. What I can say for sure is that we certainly did not produce damage, but I have the feeling that we did not close the gap and we remain at the mercy of search engines and their algorithms, of which we know very little. 

That is not to say that we are surrendering. On the contrary, we recognize that competition became even stiffer with the advent of artificial intelligence, which is sending another shockwave through global academic community. Now that the elephant of AI is in the room, opinions are divided. Some are trying to slow it down or are simply ignoring it. The others are attempting to ride the wave and hope that they can even direct it. There are also those who think how they can profit from it, but also those genuinely concerned about its consequences. The illustration of the situation is here, prompted by me but done by AI in thirty seconds:

Belonging to this concerned batch, we recognize that it is difficult to say what impact will AI shockwave have. We are aware that it cannot be stopped, have no illusions about riding it, but think that it is safe to conclude that those who will not swim are likely to drown. The key question for us is whether AI tools can help reinserting peripheral academic production into the global scholarly exchange. Furthermore, another question is how to use it in a responsible manner, in accordance with standards of academic integrity.

So let me talk you through our first jab in that direction. The goal is to experiment on my own book in different modalities, respecting both publishing rights and principles of academic honesty, and measuring whether its visibility will increase in the 2026-2030 period as compared to 2020-2025. 

Given that conventional ways of advertising the book are by now exhausted, a starting point was the selection of tools to promote the book’s main findings. Experimenting with different modalities, we immediately excluded using paid programs. It seemed like an unfair advantage, even if the actual cost was negligible. Therefore, I used tools accessible to everybody. The starting point was to extract, optically read and transform the English summary into word by using ILovePdf program. This document was uploaded into Google’s AI regime chatbot Gemini. The first dilemma was whether to let Gemini temper with the text. Ultimately, I prompted it to lightly edit the text. Upon realizing that the changes are cosmetic and are not compromising the authorship, I documented the exchange with the chatbot. I checked, adapted and adopted this English version and used Gemini to generate translations into German, Serbian, Romanian, French and Russian. The translations were uneven, and I corrected them to a degree. In my view, an author is rarely able to produce all those translations. They would anyhow be done by translators, which justifies the use of AI in this respect naturally as long as he has the last word. Still, I did not put any of this under my name, as I was at best a corrector, not the author of the translations. There were further temptations. For instance, Gemini asked me if I want to see a table of my key terms juxtaposed in all those languages and displayed in a comprehensive way. I could have done it on my own, but it would take me a while, whereas Gemini did it within seconds. Was that simple cutting corners or actual cheating? In my view, not recognizing that the machine not only did the work, but actually proposed to do it would be academically dishonest. Even more dangerous is the rabbit hole which opens with further questions that Gemini asked me, offering to put the names of politicians and academics next to specific concepts. It also offered to illustrate my book. That’s where i drew the line, because the machine started prompting me, instead of the other way around. It is an absolute imperative to document and stay in control of this process, both in order to retain the human authorship over the text and to avoid hallucinations which appear once machine’s linguistic model interfered with your own writing.

The same goes for the illustrations for the cover page, whose actual outlook could be inspired by the author, but the execution belongs to the illustrator. In that respect, there is no reason not to cooperate with the AI on the task, however with caution. Based on my previous input, Gemini offered this visualization for the book cover:

Observe that parts of the cover text. For a moment I thought the text might be in Turkish or Albanian, as I do not speak those languages. I checked with Gemini which claimed that it is in Serbian and German. Only when I pushed back, it admitted that it just spurted non-sensical words, akin to ipsum lorem tradition. It also could not explain why it turned me into Vladimir Iilic, nor tell me if this photo exists and is it protected by copyrights or available through Creative Commons.  To me that illustrates very well the hidden dangers of non-reflective usage of AI. Hence, I discontinued my cooperation with Gemini, asking it for recommendation of a free software for generating images. It recommended Perchance AI image generator, which was indeed more able to respond to a prompt and enter into a constructive dialogue, creating variants of abstract illustrations which were not under copyrights. I would have taken them into consideration. Although i personally preferred my own choice for a cover, which was a painting of Petar Omcikus called the Tailors of History, i had to admit that i could live with this one too, especially as the program was offering different variations.

Now, where the experiment went sticky was when I prompted academia.edu to produce a comic based on the summary, which looked like this:Although it did catch the gist of an argument, i had a feeling that the authorship is genuinely lost at the expense of trivialization and simplification in order to draw an attention. This is probably an area where one should stop using generative AI or would consult with a comic maker beforehand and attempt to influence the text, as well as the overall tone and message. My recommendation would also be to ascribe the authorship to a program, or to a program and the human, with extensive archiving of the process. The rule of the thumb would be that one should avoid (1) plagiarising the machine, which obviously is on the level of a sophisticated creation and (2) avoiding being prompted by the machine and losing control over the creative process of writing and thinking about your work.  

Similar principles could apply for the process of propagation of your work. Most generative AI’s are very good in making summaries, blurbs and other short versions of your text. That works exceptionally well if you forbid a machine to use only the text you uploaded, and basically to cut its own access to the outer world. However, since there are no guarantees that it will do that, thorough checking of those blurbs could save one from a considerable embarrassment.

However, at this point much more sophisticated propaganda can be created then blurbs. Take available for free program NotebookLM owned by Google. It offers a wide range of forms digesting your input. For instance, it offered convincing blurbs, followed by presentations and AI generated podcasts. This is where we enter a very tricky field. The program was very good in making a blurb. It also had no problem in making presentations in different languages (see slides in English and slides followed by narration in German). However, this is where it started making its own choices. On both occasions it chose not to summarize the entire book. English slides were ending with revolutionary purges, and German narration stopped even before, at the end of 19th century. That was already a complete distortion. What it created was nice, interesting, informed by some of my ideas, but it clearly took a life of its own. Not only that it would be immoral to sign these presentations with my name. That is especially the case since NotebookLM in its free form is nonnegotiable, you cannot really change what it made.

The same problem appeared again in the most appealing option of NotebookLM, which can generate podcasts in which two ‘people’talk about your book. The way it consumes, digests and spills back the material you feed it with is very impressive at a first glance. You can decide whether you want 5, 20 or 40 minutes long podcast. However, upon a closer inspection, it also has a tendency to choose its own ending and also to create parallels and metaphors of its own, certainly through using the material other than the one you fed it with. It does so in such a flattering way that it is easy to go with a flow and even to forget your original point. In my case it created an interesting podcast which it entitled The Bloody Roots of Ethnic Cleansing. It covers a transition from religiously inspired cleansing to revolutionary cleansing, but it omits my main point.

In short, the key finding of my book was that this coinage was not invented during the 1990s Yugoslav wars. Although it did take the world by storm in that period, it had a rich prehistory. The term cleansing was used since the antiquity as a euphemism for mass atrocities. In its first iteration, it had a strong religious connotation, which was further augmented by monotheism, leading to internal cleansings of the heretics and to external cleansings of the territory from people of different faith through crusades or jihads, advancing further toward secularisation of violence through revolutionary, racial and ethnic cleansing, peaking in the Second World War when the term was actually coined in 1941, both in Romanian and Serbian language.

The only way i could make it discuss the entirety of my argument was to make NotebookLM create another podcast to emphasize the importance of eugenics and its transformation into racial cleansing in theory and practice of the Second World War, a pandemonium which produced both carnage of unseen proportions and the very coinage whose origins we were searching for. He entitled that one The Balkan Birth of Ethnic Cleansing. Taken together, these two podcasts are more or less conveying to the world what i wanted to say, in a form which is highly appealing, but is not my own. It probably advances the marketing of your book, but it also demonstrates how easy it is to forget that it was done by the machine which creates an echo chamber that leads into illusion that your ideas are widely discussed across the globe, when this is actually not the case.

In conclusion, there are many ways to augment and enhance your research through AI, but they all carry risks which have to be carefully studies in order to assess whether they are furthering our outreach or are pushing us into plagiarizing, combining our findings with the machine which is ultimately based on linguistic models, and even abdicating from the very authorship over our own work.

To be continued…

Vladimir Petrovic and his machines

 

 

Search

You are using an outdated browser which can not show modern web content.

We suggest you download Chrome or Firefox.