Skip to main content

Research Repository

Advanced Search

A data mining approach to ontology learning for automatic content-related question-answering in MOOCs.

Shatnawi, Safwan


Safwan Shatnawi


Mohamed Medhat Gaber


The advent of Massive Open Online Courses (MOOCs) allows massive volume of registrants to enrol in these MOOCs. This research aims to offer MOOCs registrants with automatic content related feedback to fulfil their cognitive needs. A framework is proposed which consists of three modules which are the subject ontology learning module, the short text classification module, and the question answering module. Unlike previous research, to identify relevant concepts for ontology learning a regular expression parser approach is used. Also, the relevant concepts are extracted from unstructured documents. To build the concept hierarchy, a frequent pattern mining approach is used which is guided by a heuristic function to ensure that sibling concepts are at the same level in the hierarchy. As this process does not require specific lexical or syntactic information, it can be applied to any subject. To validate the approach, the resulting ontology is used in a question-answering system which analyses students' content-related questions and generates answers for them. Textbook end of chapter questions/answers are used to validate the question-answering system. The resulting ontology is compared vs. the use of Text2Onto for the question-answering system, and it achieved favourable results. Finally, different indexing approaches based on a subject's ontology are investigated when classifying short text in MOOCs forum discussion data; the investigated indexing approaches are: unigram-based, concept-based and hierarchical concept indexing. The experimental results show that the ontology-based feature indexing approaches outperform the unigram-based indexing approach. Experiments are done in binary classification and multiple labels classification settings . The results are consistent and show that hierarchical concept indexing outperforms both concept-based and unigram-based indexing. The BAGGING and random forests classifiers achieved the best result among the tested classifiers.


SHATNAWI, S.M.I. 2016. A data mining approach to ontology learning for automatic content-related question-answering in MOOCs. Robert Gordon University, PhD thesis.

Thesis Type Thesis
Deposit Date Jan 23, 2017
Publicly Available Date Jan 23, 2017
Keywords Data mining; Ontology learning; Question answering system; MOOCs; Short text classification; Frequent pattern mining; Association rule mining
Public URL
Award Date Oct 31, 2016


Downloadable Citations