Which of the following is a generative classification algorithm?🔒 Đáp án trong Source VIPDBM301Đã xuất hiện trong 1 đề thi
Consider the database containing the transactions $T_1 = \{a_1, a_2, a_3\}$, $T_2 = \{a_2, a_3, a_4\}$, $T_3 = \{a_1, a_3, a_4\}$. Let mini-support (minsup) = 2. Which of the following frequent patterns is NOT closed?🔒 Đáp án trong Source VIPDBM301Đã xuất hiện trong 1 đề thi
Is word segmentation on Chinese easier than English?🔒 Đáp án trong Source VIPDBM301Đã xuất hiện trong 1 đề thi
What is TRUE about the mixture model and topic modeling?🔒 Đáp án trong Source VIPDBM301Đã xuất hiện trong 1 đề thi
If we know the support of itemset $\{a,b\}$ is 10, which of the following numbers are the possible supports of itemset $\{a,b,c\}$? Select all that apply.🔒 Đáp án trong Source VIPDBM301Đã xuất hiện trong 1 đề thi
Self-organizing maps are an example of ____.🔒 Đáp án trong Source VIPDBM301Đã xuất hiện trong 1 đề thi
True or false? In the EM algorithm, the model likelihood monotonically increases.🔒 Đáp án trong Source VIPDBM301Đã xuất hiện trong 1 đề thi
Which evaluation method is best for clustering results of a large collection of documents?🔒 Đáp án trong Source VIPDBM301Đã xuất hiện trong 1 đề thi
Consider a spatial database that consists of 1000 records. If an item A appears 200 times in the database and the rule “if A, then B” appears 100 times, what are the support and confidence for the rule “if A, then B”?🔒 Đáp án trong Source VIPDBM301Đã xuất hiện trong 1 đề thi
True or false? Assume we are using word n-grams as features to perform sentiment classification. Then higher values of $n$ will usually be less prone to overfitting (i.e., for higher values of $n$, the difference between training and testing accuracies will be smaller).🔒 Đáp án trong Source VIPDBM301Đã xuất hiện trong 1 đề thi
Which of the following issue is considered before investing in Data Mining?🔒 Đáp án trong Source VIPDBM301Đã xuất hiện trong 1 đề thi
Let $w$ be a word and $X_w$ be a binary random variable that indicates whether $w$ appears in a text document in the corpus. Assume that the probability $P(X_w=1)$ is estimated by $\dfrac{\operatorname{Count}(w)}{N}$, where $\operatorname{Count}(w)$ is the number of documents that $w$ appears in and $N$ is the total number of documents in the corpus. \nYou are given that "the" is a very frequent word that appears in 99% of the documents and that "photon" is a very rare word that occurs in 1% of the documents. Which word has a higher entropy?🔒 Đáp án trong Source VIPDBM301Đã xuất hiện trong 1 đề thi