Document details

Ontologies for reusing data cleaning knowledge

Author(s): Almeida, Ricardo ; Oliveira, Paulo ; Braga, Luís ; Barroso, João

Date: 2012

Persistent ID: http://hdl.handle.net/10400.22/1567

Origin: Repositório Científico do Instituto Politécnico do Porto

Subject(s): Data cleaning; Ontologies; OWL; Data quality; Data cleaning; Data cleaning; Ontologies; Ontologies; OWL; OWL; Data quality; Data quality


Description

The emergence of new business models, namely, the establishment of partnerships between organizations, the chance that companies have of adding existing data on the web, especially in the semantic web, to their information, led to the emphasis on some problems existing in databases, particularly related to data quality. Poor data can result in loss of competitiveness of the organizations holding these data, and may even lead to their disappearance, since many of their decision-making processes are based on these data. For this reason, data cleaning is essential. Current approaches to solve these problems are closely linked to database schemas and specific domains. In order that data cleaning can be used in different repositories, it is necessary for computer systems to understand these data, i.e., an associated semantic is needed. The solution presented in this paper includes the use of ontologies: (i) for the specification of data cleaning operations and, (ii) as a way of solving the semantic heterogeneity problems of data stored in different sources. With data cleaning operations defined at a conceptual level and existing mappings between domain ontologies and an ontology that results from a database, they may be instantiated and proposed to the expert/specialist to be executed over that database, thus enabling their interoperability.

Document Type Conference object
Language English
Contributor(s) REPOSITÓRIO P.PORTO
facebook logo  linkedin logo  twitter logo 
mendeley logo

Related documents

No related documents