Giovanni Semeraro

Also published as: G. Semeraro


2023

pdf bib
XL-LEXEME: WiC Pretrained Model for Cross-Lingual LEXical sEMantic changE
Pierluigi Cassotti | Lucia Siciliani | Marco DeGemmis | Giovanni Semeraro | Pierpaolo Basile
Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)

The recent introduction of large-scale datasets for the WiC (Word in Context) task enables the creation of more reliable and meaningful contextualized word embeddings.However, most of the approaches to the WiC task use cross-encoders, which prevent the possibility of deriving comparable word embeddings.In this work, we introduce XL-LEXEME, a Lexical Semantic Change Detection model.XL-LEXEME extends SBERT, highlighting the target word in the sentence. We evaluate XL-LEXEME on the multilingual benchmarks for SemEval-2020 Task 1 - Lexical Semantic Change (LSC) Detection and the RuShiftEval shared task involving five languages: English, German, Swedish, Latin, and Russian.XL-LEXEME outperforms the state-of-the-art in English, German and Swedish with statistically significant differences from the baseline results and obtains state-of-the-art performance in the RuShiftEval shared task.

2022

pdf bib
An NLP Approach for the Analysis of Global Reporting Initiative Indexes from Corporate Sustainability Reports
Marco Polignano | Nicola Bellantuono | Francesco Paolo Lagrasta | Sergio Caputo | Pierpaolo Pontrandolfo | Giovanni Semeraro
Proceedings of the First Computing Social Responsibility Workshop within the 13th Language Resources and Evaluation Conference

Sustainability reporting has become an annual requirement in many countries and for certain types of companies. Sustainability reports inform stakeholders about companies’ commitment to sustainable development and their economic, social, and environmental sustainability practices. However, the fact that norms and standards allow a certain discretion to be adopted by drafting organizations makes such reports hardly comparable in terms of layout, disclosures, key performance indicators (KPIs), and so on. In this work, we present a system based on natural language processing and information extraction techniques to retrieve relevant information from sustainability reports, compliant with the Global Reporting Initiative Standards, written in Italian and English language. Specifically, the system is able to identify references to the various sustainability topics discussed by the reports: on which page of the document those references have been found, the context of each reference, and if it is mentioned positively or negatively. The output of the system has been then evaluated against a ground truth obtained through a manual annotation process on 134 reports. Experimental outcomes highlight the affordability of the approach for improving sustainability disclosures, accessibility, and transparency, thus empowering stakeholders to conduct further analysis and considerations.

pdf bib
swapUNIBA@FinTOC2022: Fine-tuning Pre-trained Document Image Analysis Model for Title Detection on the Financial Domain
Pierluigi Cassotti | Cataldo Musto | Marco DeGemmis | Georgios Lekkas | Giovanni Semeraro
Proceedings of the 4th Financial Narrative Processing Workshop @LREC2022

In this paper, we introduce the results of our submitted system to the FinTOC 2022 task. We address the task using a two-stage process: first, we detect titles using Document Image Analysis, then we train a supervised model for the hierarchical level prediction. We perform Document Image Analysis using a pre-trained Faster R-CNN on the PublyaNet dataset. We fine-tuned the model on the FinTOC 2022 training set. We extract orthographic and layout features from detected titles and use them to train a Random Forest model to predict the title level. The proposed system ranked #1 on both Title Detection and the Table of Content extraction tasks for Spanish. The system ranked #3 on both the two subtasks for English and French.

2019

pdf bib
Diachronic Analysis of Entities by Exploiting Wikipedia Page revisions
Pierpaolo Basile | Annalina Caputo | Seamus Lawless | Giovanni Semeraro
Proceedings of the International Conference on Recent Advances in Natural Language Processing (RANLP 2019)

In the last few years, the increasing availability of large corpora spanning several time periods has opened new opportunities for the diachronic analysis of language. This type of analysis can bring to the light not only linguistic phenomena related to the shift of word meanings over time, but it can also be used to study the impact that societal and cultural trends have on this language change. This paper introduces a new resource for performing the diachronic analysis of named entities built upon Wikipedia page revisions. This resource enables the analysis over time of changes in the relations between entities (concepts), surface forms (words), and the contexts surrounding entities and surface forms, by analysing the whole history of Wikipedia internal links. We provide some useful use cases that prove the impact of this resource on diachronic studies and delineate some possible future usage.

pdf bib
SWAP at SemEval-2019 Task 3: Emotion detection in conversations through Tweets, CNN and LSTM deep neural networks
Marco Polignano | Marco de Gemmis | Giovanni Semeraro
Proceedings of the 13th International Workshop on Semantic Evaluation

Emotion detection from user-generated contents is growing in importance in the area of natural language processing. The approach we proposed for the EmoContext task is based on the combination of a CNN and an LSTM using a concatenation of word embeddings. A stack of convolutional neural networks (CNN) is used for capturing the hierarchical hidden relations among embedding features. Meanwhile, a long short-term memory network (LSTM) is used for capturing information shared among words of the sentence. Each conversation has been formalized as a list of word embeddings, in particular during experimental runs pre-trained Glove and Google word embeddings have been evaluated. Surface lexical features have been also considered, but they have been demonstrated to be not usefully for the classification in this specific task. The final system configuration achieved a micro F1 score of 0.7089. The python code of the system is fully available at https://github.com/marcopoli/EmoContext2019

2017

pdf bib
Centroid-based Text Summarization through Compositionality of Word Embeddings
Gaetano Rossiello | Pierpaolo Basile | Giovanni Semeraro
Proceedings of the MultiLing 2017 Workshop on Summarization and Summary Evaluation Across Source Types and Genres

The textual similarity is a crucial aspect for many extractive text summarization methods. A bag-of-words representation does not allow to grasp the semantic relationships between concepts when comparing strongly related sentences with no words in common. To overcome this issue, in this paper we propose a centroid-based method for text summarization that exploits the compositional capabilities of word embeddings. The evaluations on multi-document and multilingual datasets prove the effectiveness of the continuous vector representation of words compared to the bag-of-words model. Despite its simplicity, our method achieves good performance even in comparison to more complex deep learning models. Our method is unsupervised and it can be adopted in other summarization tasks.

2015

pdf bib
UNIBA: Combining Distributional Semantic Models and Sense Distribution for Multilingual All-Words Sense Disambiguation and Entity Linking
Pierpaolo Basile | Annalina Caputo | Giovanni Semeraro
Proceedings of the 9th International Workshop on Semantic Evaluation (SemEval 2015)

2014

pdf bib
UNIBA: Combining Distributional Semantic Models and Word Sense Disambiguation for Textual Similarity
Pierpaolo Basile | Annalina Caputo | Giovanni Semeraro
Proceedings of the 8th International Workshop on Semantic Evaluation (SemEval 2014)

pdf bib
An Enhanced Lesk Word Sense Disambiguation Algorithm through a Distributional Semantic Model
Pierpaolo Basile | Annalina Caputo | Giovanni Semeraro
Proceedings of COLING 2014, the 25th International Conference on Computational Linguistics: Technical Papers

2013

pdf bib
UNIBA-CORE: Combining Strategies for Semantic Textual Similarity
Annalina Caputo | Pierpaolo Basile | Giovanni Semeraro
Second Joint Conference on Lexical and Computational Semantics (*SEM), Volume 1: Proceedings of the Main Conference and the Shared Task: Semantic Textual Similarity

2012

pdf bib
UNIBA: Distributional Semantics for Textual Similarity
Annalina Caputo | Pierpaolo Basile | Giovanni Semeraro
*SEM 2012: The First Joint Conference on Lexical and Computational Semantics – Volume 1: Proceedings of the main conference and the shared task, and Volume 2: Proceedings of the Sixth International Workshop on Semantic Evaluation (SemEval 2012)

2011

pdf bib
Encoding syntactic dependencies by vector permutation
Pierpaolo Basile | Annalina Caputo | Giovanni Semeraro
Proceedings of the GEMS 2011 Workshop on GEometrical Models of Natural Language Semantics

2010

pdf bib
UBA: Using Automatic Translation and Wikipedia for Cross-Lingual Lexical Substitution
Pierpaolo Basile | Giovanni Semeraro
Proceedings of the 5th International Workshop on Semantic Evaluation

2008

pdf bib
Combining Knowledge-based Methods and Supervised Learning for Effective Italian Word Sense Disambiguation
Pierpaolo Basile | Marco de Gemmis | Pasquale Lops | Giovanni Semeraro
Semantics in Text Processing. STEP 2008 Conference Proceedings

2007

pdf bib
UNIBA: JIGSAW algorithm for Word Sense Disambiguation
Pierpaolo Basile | Marco de Gemmis | Anna Lisa Gentile | Pasquale Lops | Giovanni Semeraro
Proceedings of the Fourth International Workshop on Semantic Evaluations (SemEval-2007)

2000

pdf bib
A Semi-automatic System for Conceptual Annotation, its Application to Resource Construction and Evaluation
W.J. Black | J. McNaught | G.P. Zarri | A. Persidis | A. Brasher | L. Gilardoni | E. Bertino | G. Semeraro | P. Leo
Proceedings of the Second International Conference on Language Resources and Evaluation (LREC’00)

pdf bib
Learning from Parsed Sentences with INTHELEX
F. Esposito | S. Ferilli | N. Fanizzi | G. Semeraro
Fourth Conference on Computational Natural Language Learning and the Second Learning Language in Logic Workshop