Utilize este identificador para referenciar este registo: https://hdl.handle.net/10316/100540
Título: Assessment of SQL and NoSQL Systems to Store and Mine COVID-19 Data
Autor: Antas, João
Silva, Rodrigo Rocha 
Bernardino, Jorge 
Palavras-chave: big data; COVID-19; Data Mining; SQL and NoSQL databases
Data: 2022
Título da revista, periódico, livro ou evento: Computers
Volume: 11
Número: 2
Resumo: COVID-19 has provoked enormous negative impacts on human lives and the world economy. In order to help in the fight against this pandemic, this study evaluates different databases’ systems and selects the most suitable for storing, handling, and mining COVID-19 data. We evaluate different SQL and NoSQL database systems using the following metrics: query runtime, memory used, CPU used, and storage size. The databases systems assessed were Microsoft SQL Server, MongoDB, and Cassandra. We also evaluate Data Mining algorithms, including Decision Trees, Random Forest, Naive Bayes, and Logistic Regression using Orange Data Mining software data classification tests. Classification tests were performed using cross-validation in a table with about 3 M records, including COVID-19 exams with patients’ symptoms. The Random Forest algorithm has obtained the best average accuracy, recall, precision, and F1 Score in the COVID-19 predictive model performed in the mining stage. In performance evaluation, MongoDB has presented the best results for almost all tests with a large data volume.
URI: https://hdl.handle.net/10316/100540
ISSN: 2073-431X
DOI: 10.3390/computers11020029
Direitos: openAccess
Aparece nas coleções:I&D CISUC - Artigos em Revistas Internacionais

Ficheiros deste registo:
Mostrar registo em formato completo

Citações SCOPUSTM   

6
Visto em 15/abr/2024

Citações WEB OF SCIENCETM

3
Visto em 2/abr/2024

Visualizações de página

66
Visto em 16/abr/2024

Downloads

46
Visto em 16/abr/2024

Google ScholarTM

Verificar

Altmetric

Altmetric


Este registo está protegido por Licença Creative Commons Creative Commons