Enhanced geographically typed semantic schema matching

Jeffrey Partyka; Pallabi Parveen; Latifur Khan; B. Thuraisingham; Shashi Shekhar

doi:10.1016/j.websem.2010.11.002

Enhanced geographically typed semantic schema matching

Jeffrey Partyka, Pallabi Parveen, Latifur Khan, B. Thuraisingham, Shashi Shekhar

Computer Science and Engineering

Research output: Contribution to journal › Article › peer-review

12 Scopus citations

Abstract

Resolving semantic heterogeneity across distinct data sources remains a highly relevant problem in the GIS domain requiring innovative solutions. Our approach, called GSim, semantically aligns tables from respective GIS databases by first choosing attributes for comparison. We then examine their instances and calculate a similarity value between them called entropy-based distribution (EBD)¹ by combining two separate methods. Our primary method discerns the geographic types from instances of compared attributes. If successful, EBD is calculated using only this method. GSim further facilitates geographic type matching by using latlong values to further disambiguate between multiple types of a given instance and applying attribute weighting to quantify the uniqueness of mapped attributes. If geographic type matching is not possible, we then apply a generic schema matching method, independent of the knowledge domain, which employs normalized Google distance. We show the effectiveness of our approach over the traditional approaches across multi-jurisdictional datasets by generating impressive results.

Original language	English (US)
Pages (from-to)	52-70
Number of pages	19
Journal	Journal of Web Semantics
Volume	9
Issue number	1
DOIs	https://doi.org/10.1016/j.websem.2010.11.002
State	Published - Mar 1 2011

Keywords

GIS
Gazetteer
Geocoding
Geosemantics
Geotypes
Schema

Access

10.1016/j.websem.2010.11.002

OpenUrl availability

Full text

Cite this

@article{2c51038886d84692bc6659808f056d8f,

title = "Enhanced geographically typed semantic schema matching",

abstract = "Resolving semantic heterogeneity across distinct data sources remains a highly relevant problem in the GIS domain requiring innovative solutions. Our approach, called GSim, semantically aligns tables from respective GIS databases by first choosing attributes for comparison. We then examine their instances and calculate a similarity value between them called entropy-based distribution (EBD)1 by combining two separate methods. Our primary method discerns the geographic types from instances of compared attributes. If successful, EBD is calculated using only this method. GSim further facilitates geographic type matching by using latlong values to further disambiguate between multiple types of a given instance and applying attribute weighting to quantify the uniqueness of mapped attributes. If geographic type matching is not possible, we then apply a generic schema matching method, independent of the knowledge domain, which employs normalized Google distance. We show the effectiveness of our approach over the traditional approaches across multi-jurisdictional datasets by generating impressive results.",

keywords = "GIS, Gazetteer, Geocoding, Geosemantics, Geotypes, Schema",

author = "Jeffrey Partyka and Pallabi Parveen and Latifur Khan and B. Thuraisingham and Shashi Shekhar",

year = "2011",

month = mar,

day = "1",

doi = "10.1016/j.websem.2010.11.002",

language = "English (US)",

volume = "9",

pages = "52--70",

journal = "Journal of Web Semantics",

issn = "1570-8268",

publisher = "Elsevier",

number = "1",

}

TY - JOUR

T1 - Enhanced geographically typed semantic schema matching

AU - Partyka, Jeffrey

AU - Parveen, Pallabi

AU - Khan, Latifur

AU - Thuraisingham, B.

AU - Shekhar, Shashi

PY - 2011/3/1

Y1 - 2011/3/1

N2 - Resolving semantic heterogeneity across distinct data sources remains a highly relevant problem in the GIS domain requiring innovative solutions. Our approach, called GSim, semantically aligns tables from respective GIS databases by first choosing attributes for comparison. We then examine their instances and calculate a similarity value between them called entropy-based distribution (EBD)1 by combining two separate methods. Our primary method discerns the geographic types from instances of compared attributes. If successful, EBD is calculated using only this method. GSim further facilitates geographic type matching by using latlong values to further disambiguate between multiple types of a given instance and applying attribute weighting to quantify the uniqueness of mapped attributes. If geographic type matching is not possible, we then apply a generic schema matching method, independent of the knowledge domain, which employs normalized Google distance. We show the effectiveness of our approach over the traditional approaches across multi-jurisdictional datasets by generating impressive results.

AB - Resolving semantic heterogeneity across distinct data sources remains a highly relevant problem in the GIS domain requiring innovative solutions. Our approach, called GSim, semantically aligns tables from respective GIS databases by first choosing attributes for comparison. We then examine their instances and calculate a similarity value between them called entropy-based distribution (EBD)1 by combining two separate methods. Our primary method discerns the geographic types from instances of compared attributes. If successful, EBD is calculated using only this method. GSim further facilitates geographic type matching by using latlong values to further disambiguate between multiple types of a given instance and applying attribute weighting to quantify the uniqueness of mapped attributes. If geographic type matching is not possible, we then apply a generic schema matching method, independent of the knowledge domain, which employs normalized Google distance. We show the effectiveness of our approach over the traditional approaches across multi-jurisdictional datasets by generating impressive results.

KW - GIS

KW - Gazetteer

KW - Geocoding

KW - Geosemantics

KW - Geotypes

KW - Schema

UR - http://www.scopus.com/inward/record.url?scp=79951681438&partnerID=8YFLogxK

UR - http://www.scopus.com/inward/citedby.url?scp=79951681438&partnerID=8YFLogxK

U2 - 10.1016/j.websem.2010.11.002

DO - 10.1016/j.websem.2010.11.002

M3 - Article

AN - SCOPUS:79951681438

SN - 1570-8268

VL - 9

SP - 52

EP - 70

JO - Journal of Web Semantics

JF - Journal of Web Semantics

IS - 1

ER -

Enhanced geographically typed semantic schema matching

Abstract

Keywords

Access

OpenUrl availability

Other files and links

Fingerprint

Cite this