CMI Data Science
    Previous Year Papers
    Verified Solutions Included
    CMI Data Science 2022 Question Paper with Solutions: 22 Questions, Answer Key & Section-wise Analysis

    CMI Data Science 2022 previous year paper: 22 questions with answer key and detailed solutions, section-wise breakdown and free sample questions.

    22 Qs

    Total Questions

    51 Marks

    Total Marks

    1.7850000000000001 Mins

    Duration

    +3 / -1 / 0

    Marking Scheme

    Section-wise Paper Structure

    School Level Mathematics

    10 Qs

    45% of total marks

    Discrete Mathematics

    6 Qs

    27% of total marks

    Probability Theory

    5 Qs

    23% of total marks

    Programming

    1 Qs

    5% of total marks

    Free Solved Questions with Step-by-Step Solutions

    Authentic examination problems with detailed derivations and answer keys.

    Question 1
    2022 PYQ
    Level 3: Exam Standard

    Which of the following statements is/are true?

    Question 2
    2022 PYQ
    Level 3: Exam Standard

    A matrix is said to be symmetric if . Which of the following is/are true? Let be matrices.

    Question 3
    2022 PYQ
    Level 3: Exam Standard

    Which of the following statements is/are true?

    Question 4
    2022 PYQ
    Level 3: Exam Standard

    In order to select a debating team to represent a school, 7 students from class XII and 13 students from class XI were shortlisted and were undergoing trials. The coach had to select a team of 5 students, out of which at least two students should be from each class. One out of the 5 students was to be named as team leader, who was to be from class XII. Two teams with the same members but different leaders are considered to be two different teams. The number of different teams the coach can select is

    Question 5
    2022 PYQ
    Level 4: Challenger
    Words are formed using the characters 0 and 1. The length of a word is the number of characters in it. We say there is a path from word to word if starting from the word you can get the word by applying the following sequence of transformation rules finitely many times in any order.
    • Replacing an occurrence of the string 101 by a 0.
    • Replacing an occurrence of the string 010 by a 1.
    • Replacing a 0 by a 101.
    • Replacing a 1 by a 010.
    If there is a path from word to word we say and are equivalent or that they are in the same equivalence class. For example the four letter word 1011 is equivalent to the two letter word 01, since the prefix 101 in 1011 can be replaced by a 0 to get 01.
    We say has a shorter description if there is a word of shorter length equivalent to it.
    (a) State true or false: for any word there is a unique shortest word in its equivalence class.
    (b) How many three letter words are there in this language which have no shorter descriptions and which are all in different equivalence classes?
    Question 6
    2022 PYQ
    Level 3: Exam Standard
    A relation on the set is defined by reading the columns of the following table from top to bottom. If a column in the table reads it means is related to in . If a column in the table reads it means is not related to .
    abcdabcdabcdabcdaaaabbbbccccdddd0010100011010010
    For instance, from the fifth column we have, , and from the second we have .
    Another relation on the set is defined as: for any , the pair is in if and only if there exists such that both and hold.
    Which of the following pairs are in ?
    Question 7
    2022 PYQ
    Level 3: Exam Standard

    Let be a Binomial random variable with the probability mass function

    Let be a random variable defined as:

    Let and denote the expectation and the variance of a random variable . Which of the following statements is/are true?

    Question 8
    2022 PYQ
    Level 2: Moderate
    Common Description: Description for the next four questions:
    Dasholytics Inc runs an online analytics dashboard. The company employs three machines—Server 1, Server 2, and Server 3—to serve the data for the dashboard. At any point in time one of these three machines has the job of serving the data, and the other two are kept in standby mode. Dasholytics uses a server scheduler (software) that decides which of the three machines should serve the data at any given point of time. The scheduler switches between data servers without disrupting the data feed to the dashboard.
    The machine which is serving the data sometimes fails to do so; in this case an alert is sent out to the Dasholytics team and they investigate and fix the problem. The time for which a machine fails to serve data is accounted as service outage caused by that machine. For technical reasons the server scheduler is deactivated when there is such an outage: the outage gets over only when the issue with the server is fixed and it is put back online.
    The figures below describe the server usage and outage statistics as compiled over the last one year (365 days). Please use this information to answer the questions that follow.
    Server 1 Server 2 Server 3 0 1 2 3 4 5 Outage Percentage (%) Fig. (a) Server 1: 40% Server 3: 30% Server 2: 30% Fig. (b)
    Figure 1: Figure (a) shows the total service outage caused by each server, as a percentage of the total time that it served data. Figure (b) shows the total time that each server served data, as a percentage of the total duration over which this information was collected. Express the total service outage caused by all the three servers, as a percentage of the total duration over which this information was collected.
    Question 9
    2022 PYQ
    Level 2: Moderate
    Common Description: Description for the next four questions:
    Dasholytics Inc runs an online analytics dashboard. The company employs three machines—Server 1, Server 2, and Server 3—to serve the data for the dashboard. At any point in time one of these three machines has the job of serving the data, and the other two are kept in standby mode. Dasholytics uses a server scheduler (software) that decides which of the three machines should serve the data at any given point of time. The scheduler switches between data servers without disrupting the data feed to the dashboard.
    The machine which is serving the data sometimes fails to do so; in this case an alert is sent out to the Dasholytics team and they investigate and fix the problem. The time for which a machine fails to serve data is accounted as service outage caused by that machine. For technical reasons the server scheduler is deactivated when there is such an outage: the outage gets over only when the issue with the server is fixed and it is put back online.
    The figures below describe the server usage and outage statistics as compiled over the last one year (365 days). Please use this information to answer the questions that follow.
    Server 1 Server 2 Server 3 0 1 2 3 4 5 Outage Percentage (%) Fig. (a) Server 1: 40% Server 3: 30% Server 2: 30% Fig. (b)
    Figure 1: Figure (a) shows the total service outage caused by each server, as a percentage of the total time that it served data. Figure (b) shows the total time that each server served data, as a percentage of the total duration over which this information was collected. What is the expected number of hours of service outage in an year (365 days)?
    Question 10
    2022 PYQ
    Level 3: Exam Standard
    Consider the following code, in which A is an array indexed from 0 and n is the number of elements in A. function foo(A,n) { L = 0; R = n - 1; while (L <= R) { i = ceil((L + R)/2); if (A[i] < i) { L = i + 1; } else { if (A[i] > i) { R = i - 1; } else { return(i); } } } return(-1); } Here, ceil(x) returns the smallest integer bigger than or equal to the number x.
    If , what will foo(A, 10) return?

    Unlock All 22 Questions in Real Examination Mode

    Practice with the authentic timer, on-screen calculator, instant percentile ranking, and section-wise analytics.

    More CMI Data Science Previous Year Papers

    Free preview ends here

    Login to view the complete paper and solutions

    Creating an account is free. You get the rest of this chapter, step-by-step solutions, and a study plan built around the topics you are actually weak at.

    Why MastersUp

    Personalised first. High quality throughout.

    Most platforms hand everyone the same content. Here the content moves with your performance, topic by topic.

    Built around you, not around a syllabus PDF

    Every answer you give moves your topic-level intelligence rate. The next question, the next revision card and tomorrow's plan all change with it.

    Revision that hits your weak spots

    We only revise topics you have actually attempted and are still below the safe bar on — never the same chapter on repeat.

    Questions calibrated to the real exam

    Each question carries a measured toughness. You are served a rung above your current level, so practice keeps stretching you.

    Notes written for recall, not for volume

    Full lesson cards for first study, curated short-note cards for the last mile — with derivations, traps and exam patterns marked.

    One place for everything

    Notes, chapter practice, previous-year questions, test series and full-length papers — all feeding one picture of your preparation.

    Honest progress

    No vanity streaks. Progress here means chapters mastered and accuracy that held up on harder questions.

    Unlock the whole course

    Full notes and short notes, the complete question bank with worked solutions, mock tests, full-length papers, and an adaptive plan that rebuilds itself as you improve.

    CMI Data Science 2022 Question Paper with Solutions: 22 Questions, Answer Key & Section-wise Analysis

    CMI Data Science 2022 previous year paper: 22 questions with answer key and detailed solutions, section-wise breakdown and free sample questions.

    Paper breakdown

    22 questions · 51 marks · 1.7850000000000001 minutes. School Level Mathematics: 10 · Discrete Mathematics: 6 · Probability Theory: 5 · Programming: 1

    Free sample questions from CMI Data Science 2022 Question Paper

    Question 1 · School Level Mathematics · 2022 MSQ

    Which of the following statements is/are true?

    1. A.

      For any real number with , .

    2. B.

      Let be the positive square root of and let be any real number. Then

    3. C.

    4. D.

      The function defined as

    Correct Answer:

    ["B","D"]

    Step-by-Step Solution

    Key idea: This is a multi-concept statement-truth question testing geometric series convergence, Euler's formula, telescoping series, and probability density function validation.

    Step 1: Analyze Option A

    The statement claims: For |r| > 1, lim_{n→∞} Σ_{i=1}^{n} 17r^n = 17/(1-r)

    First, note the sum notation: Σ_{i=1}^{n} 17r^n. The index is i but the term 17r^n doesn't depend on i. This equals n · 17r^n.

    Even if this is a typo and should be Σ 17r^i, for |r| > 1, the terms grow without bound and the series DIVERGES. The formula 17/(1-r) only applies when |r| < 1.

    Therefore, Option A is FALSE.

    Step 2: Analyze Option B

    We need to verify: tan x = (e^{ix} - e^{-ix})/(i(e^{ix} + e^{-ix}))

    From Euler's formula:

    • e^{ix} = cos x + i sin x
    • e^{-ix} = cos x - i sin x

    Therefore:

    • e^{ix} - e^{-ix} = 2i sin x
    • e^{ix} + e^{-ix} = 2 cos x

    Substituting:

    (e^{ix} - e^{-ix})/(i(e^{ix} + e^{-ix})) = (2i sin x)/(i · 2 cos x) = sin x/cos x = tan x

    This matches the given expression. Option B is TRUE.

    Step 3: Analyze Option C

    The series is: 1/(1·2) + 1/(2·3) + 1/(3·4) + ...

    Using partial fractions: 1/(n(n+1)) = 1/n - 1/(n+1)

    The partial sum up to n terms is:

    S_n = (1 - 1/2) + (1/2 - 1/3) + (1/3 - 1/4) + ... + (1/n - 1/(n+1))

    S_n = 1 - 1/(n+1) [telescoping cancellation]

    As n → ∞: S_∞ = lim_{n→∞} (1 - 1/(n+1)) = 1

    The infinite sum EQUALS exactly 1, not less than 1. The statement says the sum is < 1, which is FALSE.

    Step 4: Analyze Option D

    The function is defined as:

    f(x) = 3cos(3x) for 0 < x < π/6

    f(x) = 0 otherwise

    For f(x) to be a valid probability density function (PDF), we need:

    1. f(x) ≥ 0 for all x
    2. ∫_{-∞}^{∞} f(x) dx = 1

    Checking condition 1: For 0 < x < π/6, we have 0 < 3x < π/2, so cos(3x) > 0. Thus 3cos(3x) > 0. ✓

    Checking condition 2:

    ∫_{-∞}^{∞} f(x) dx = ∫_{0}^{π/6} 3cos(3x) dx

    = [sin(3x)]_{0}^{π/6}

    = sin(π/2) - sin(0)

    = 1 - 0

    = 1 ✓

    Both conditions are satisfied. Option D is TRUE.

    Answer: B, D

    Question 2 · School Level Mathematics · 2022 MSQ

    A matrix is said to be symmetric if . Which of the following is/are true? Let be matrices.

    1. A.

      If is symmetric and invertible, then is also symmetric and invertible.

    2. B.

      If and are symmetric, then is also symmetric.

    3. C.

      If and are invertible, then is also invertible.

    4. D.

      If and are symmetric, then is also symmetric.

    Correct Answer:

    ["A","C","D"]

    Step-by-Step Solution

    Key idea: This is a closure properties question, recognisable because it tests whether symmetry and invertibility are preserved under matrix addition, multiplication, and inversion.

    Step 1: Analyze Option A. "If is symmetric and invertible, then is also symmetric and invertible."

    We know . Since is symmetric, , so . Thus, is symmetric. It is invertible by definition. True.

    Step 2: Analyze Option B. "If and are symmetric, then is also symmetric."

    Check the transpose: . For to be symmetric, we need . Since matrix multiplication is not generally commutative, is not necessarily symmetric. False.

    Step 3: Analyze Option C. "If and are invertible, then is also invertible."

    The determinant of a product is the product of determinants: . Since and are invertible, and , so . Thus, is invertible. True.

    Step 4: Analyze Option D. "If and are symmetric, then is also symmetric."

    Check the transpose: . Thus, is symmetric. True.

    Answer: A, C, D

    Question 3 · School Level Mathematics · 2022 MSQ

    Which of the following statements is/are true?

    1. A.

      Let be a matrix with . Then there is a nonzero solution with only if has full row rank.

    2. B.

      Let be a matrix with . There is a nonzero solution with .

    3. C.

      The row rank of an matrix is equal to its column rank only when .

    4. D.

      Let be an matrix. Suppose , where has size and has size . The rank of is less than or equal to .

    Correct Answer:

    ["D"]

    Step-by-Step Solution

    Key idea: This is a rank and null space question, recognisable because it tests the relationship between matrix dimensions, rank, nullity, and matrix factorization.

    Step 1: Analyze Option A. "Let be with . Nonzero solution only if has full row rank."

    By the Rank-Nullity Theorem, . Since , we have . Thus, a nonzero solution always exists, regardless of whether has full row rank or not. The "only if" condition makes this statement False.

    Step 2: Analyze Option B. "Let be with . There is a nonzero solution ."

    Here, . If has full column rank (), then , meaning the only solution is . Thus, a nonzero solution is not guaranteed. False.

    Step 3: Analyze Option C. "The row rank of an matrix is equal to its column rank only when ."

    A fundamental theorem of linear algebra states that row rank always equals column rank for any matrix, regardless of its dimensions. False.

    Step 4: Analyze Option D. "Let , where is and is . The rank of is less than or equal to ."

    By the rank inequality for matrix products, . Since has columns, . Since has rows, . Therefore, . True.

    Answer: D

    Question 4 · Discrete Mathematics · 2022 MSQ

    In order to select a debating team to represent a school, 7 students from class XII and 13 students from class XI were shortlisted and were undergoing trials. The coach had to select a team of 5 students, out of which at least two students should be from each class. One out of the 5 students was to be named as team leader, who was to be from class XII. Two teams with the same members but different leaders are considered to be two different teams. The number of different teams the coach can select is

    1. A.

    2. B.

    3. C.

    4. D.

    Correct Answer:

    ["C"]

    Step-by-Step Solution

    Key idea: This is a committee selection with class constraints and a leader requirement, identical in structure to P2. The method is: split into valid composition cases, count member selections, multiply by eligible leaders per case, then sum.

    Step 1: Identify valid compositions.

    Team of 5, at least 2 from each class:

    • Case A: 2 from XII, 3 from XI.
    • Case B: 3 from XII, 2 from XI.

    Step 2: Case A count.

    Members: .

    Leader must be from XII; there are 2 such students in this team.

    Case A: .

    Step 3: Case B count.

    Members: .

    Leader must be from XII; there are 3 such students.

    Case B: .

    Step 4: Total.

    .

    This matches option C.

    Alternative approach (leader-first): Choose the leader first from 7 class XII students. Then choose remaining 4 members ensuring at least 1 more from XII and at least 2 from XI. This gives the same result but is slightly more complex to set up.

    The trap in option D is swapping the multipliers (3 and 2), which would correspond to the leader being from class XI.

    Answer: C

    Question 5 · Discrete Mathematics · 2022 SUB
    Words are formed using the characters 0 and 1. The length of a word is the number of characters in it. We say there is a path from word to word if starting from the word you can get the word by applying the following sequence of transformation rules finitely many times in any order.
    • Replacing an occurrence of the string 101 by a 0.
    • Replacing an occurrence of the string 010 by a 1.
    • Replacing a 0 by a 101.
    • Replacing a 1 by a 010.
    If there is a path from word to word we say and are equivalent or that they are in the same equivalence class. For example the four letter word 1011 is equivalent to the two letter word 01, since the prefix 101 in 1011 can be replaced by a 0 to get 01.
    We say has a shorter description if there is a word of shorter length equivalent to it.
    (a) State true or false: for any word there is a unique shortest word in its equivalence class.
    (b) How many three letter words are there in this language which have no shorter descriptions and which are all in different equivalence classes?
    Correct Answer:

    2

    Step-by-Step Solution

    Key idea: This is a String Rewriting System with Equivalence Classes problem. The transformation rules are reversible, so equivalence is symmetric. We must find irreducible words of length 3 (no shorter description) and count how many distinct equivalence classes they form.

    Step 1: Identify all irreducible 3-letter words.

    A word has a shorter description if it contains "101" or "010" as a substring (since those can be reduced to "0" and "1" respectively).

    All 3-bit strings: 000, 001, 010, 011, 100, 101, 110, 111.

    Remove those containing "101" or "010": exclude 010 and 101.

    Irreducible set: (6 words).

    Step 2: Determine equivalence among these via expansion/reduction.

    Note: Rules are reversible, so if you can go from to using any sequence.

    Use invariant: Consider the value .

    Check how rules affect it:

    • Replace "101" (2 ones, 1 zero) with "0" (1 zero): .
    • Replace "010" (2 zeros, 1 one) with "1": .
    • Expansions are reverse, so same.

    So is invariant.

    Compute for each irreducible word:

    • 000: →
    • 001:
    • 011:
    • 100:
    • 110:
    • 111:

    So words fall into 3 groups by invariant:

    • Class 0: {000, 111}
    • Class 1: {001, 100}
    • Class 2: {011, 110}

    Step 3: Verify that within each class, words are actually equivalent.

    • 000 ↔ 111:

    Expand middle 0 in 000 → 0(101)0 = 01010

    Reduce first "010" → 110 → not 111. Alternate:

    Expand all 0s in 000: 0→101, so 000 → 101101101

    Reduce overlapping patterns… but easier: use known result that invariant is complete here.

    However, the problem states: “have no shorter descriptions and are in different equivalence classes”.

    Since the invariant partitions the set into 3 classes, and each class contains irreducible words, there are 3 classes.

    But wait—the answer is given as 2. Re-express using a better invariant.

    Alternate invariant: Consider the word as a path where 0 = +1, 1 = -1, and track cumulative sum modulo something.

    Or observe actual reductions:

    Try connecting 000 and 111:

    • 000 → expand first 0 → 10100
    • In 10100, reduce "101" → 000 → loop.
    • Expand 111: 1→010, so 111 → 010010010
    • Reduce "010" → 111 → loop.

    Now try cross-class:

    • Can 000 reach 001? Suppose yes. Then lengths mod 2 might matter.

    Count length parity:

    Reduction: "101" (len 3) → "0" (len 1): Δlen = -2

    Expansion: "0" → "101": Δlen = +2

    So length modulo 2 is invariant!

    Original words are length 3 (odd).

    Any equivalent word must have odd length.

    Shorter description would be length 1 (also odd)—allowed.

    But for irreducible words, we only care about equivalence among length-3 words.

    Now test actual reachability:

    • 000 and 111:

    Start from 0 → expand to 101 → expand first 1 to 010 → 01001

    Reduce "101" in middle? 01001 → no "101" or "010" as substring?

    "010" at start → reduce to 101 → which is reducible → reduces to 0.

    So 0 ↔ 101 ↔ 01001 ↔ 101 ↔ 0. Not helping.

    Known result in such systems: the invariant is actually the value modulo 2 of the number of 1s, or a linear combination.

    Let’s define .

    For "101" → "0": LHS: 1+2*2=5≡2; RHS:1+0=1 → not invariant.

    Simpler: notice that replacing "101" with "0" preserves the XOR of bits? No.

    Instead, manually verify connections as in the partial solution:

    Connection: 000 ↔ 110

    000 → expand middle 0 → 0 101 0 = 01010

    In 01010, "010" at start → replace with 1 → 110. So 000 ∼ 110.

    Connection: 000 ↔ 011

    011 → expand first 1 → 0 010 1 = 00101

    In 00101, "101" at end → replace with 0 → 000. So 011 ∼ 000.

    Thus, 000, 110, 011 are all equivalent.

    Similarly:

    111 → expand middle 1 → 1 010 1 = 10101

    Reduce "101" at start → 001. So 111 ∼ 001.

    111 → expand last 1 → 11 010 = 11010

    Reduce "101" in middle (positions 1-3: "101") → 100. So 111 ∼ 100.

    Thus, 111, 001, 100 are all equivalent.

    So the 6 irreducible words form exactly 2 equivalence classes:

    Class A: {000, 011, 110}

    Class B: {111, 001, 100}

    Therefore, there are 2 distinct equivalence classes of irreducible 3-letter words.

    Answer: 2

    Question 6 · Discrete Mathematics · 2022 MSQ
    A relation on the set is defined by reading the columns of the following table from top to bottom. If a column in the table reads it means is related to in . If a column in the table reads it means is not related to .
    abcdabcdabcdabcdaaaabbbbccccdddd0010100011010010
    For instance, from the fifth column we have, , and from the second we have .
    Another relation on the set is defined as: for any , the pair is in if and only if there exists such that both and hold.
    Which of the following pairs are in ?
    1. A.

    2. B.

    3. C.

    4. D.

    Correct Answer:

    ["A","C"]

    Step-by-Step Solution

    Key idea: This is a Relation Composition via Matrix/Table problem. Recognise it because relation is defined as "exists such that and ", which is precisely the definition of or .

    Step 1: Decode the table into relation .

    The table columns represent pairs . Reading the bottom row (values) against the top two rows (labels):

    Col 1:

    Col 2:

    Col 3:

    Col 4:

    Col 5:

    Col 6:

    Col 7:

    Col 8:

    Col 9:

    Col 10:

    Col 11:

    Col 12:

    Col 13:

    Col 14:

    Col 15:

    Col 16:

    So .

    Step 2: Compute .

    We need pairs where a path exists in .

    Let's trace paths of length 2 from each starting node:

    • From :

    (since )

    (since )

    • From :

    (since )

    • From :

    (since )

    (since )

    • From :

    (since )

    So .

    Step 3: Check options against .

    A: — TRUE

    B: — FALSE

    C: — TRUE

    D: — FALSE

    Answer: Options A and C are correct.

    Question 7 · Probability Theory · 2022 MSQ

    Let be a Binomial random variable with the probability mass function

    Let be a random variable defined as:

    Let and denote the expectation and the variance of a random variable . Which of the following statements is/are true?

    1. A.

      and

    2. B.

      and

    3. C.

      and

    4. D.

      and

    Correct Answer:

    ["A","C"]

    Step-by-Step Solution

    Key idea: This is an indicator-variable question. The variable is not the binomial count itself; it is a 0/1 switch that turns on only when .

    Step 1: Identify the event indicated by .

    The definition says

    So is the indicator of the event .

    Step 2: Find the probability of that event.

    Since ,

    Let

    Step 3: Use indicator properties.

    An indicator variable is Bernoulli with success probability . Therefore,

    Also, because only takes values 0 and 1,

    so

    Step 4: Compute the variance.

    For a Bernoulli indicator,

    Substituting ,

    Step 5: Match with the options.

    Option A gives the correct and correct variance.

    Option C gives the correct and correct .

    Option B incorrectly uses the mean and variance of , not .

    Option D incorrectly treats as a Bernoulli variable with parameter .

    Answer: Options A and C are true.

    Question 8 · Probability Theory · 2022 SUB
    Common Description: Description for the next four questions:
    Dasholytics Inc runs an online analytics dashboard. The company employs three machines—Server 1, Server 2, and Server 3—to serve the data for the dashboard. At any point in time one of these three machines has the job of serving the data, and the other two are kept in standby mode. Dasholytics uses a server scheduler (software) that decides which of the three machines should serve the data at any given point of time. The scheduler switches between data servers without disrupting the data feed to the dashboard.
    The machine which is serving the data sometimes fails to do so; in this case an alert is sent out to the Dasholytics team and they investigate and fix the problem. The time for which a machine fails to serve data is accounted as service outage caused by that machine. For technical reasons the server scheduler is deactivated when there is such an outage: the outage gets over only when the issue with the server is fixed and it is put back online.
    The figures below describe the server usage and outage statistics as compiled over the last one year (365 days). Please use this information to answer the questions that follow.
    Server 1 Server 2 Server 3 0 1 2 3 4 5 Outage Percentage (%) Fig. (a) Server 1: 40% Server 3: 30% Server 2: 30% Fig. (b)
    Figure 1: Figure (a) shows the total service outage caused by each server, as a percentage of the total time that it served data. Figure (b) shows the total time that each server served data, as a percentage of the total duration over which this information was collected. Express the total service outage caused by all the three servers, as a percentage of the total duration over which this information was collected.
    Correct Answer:

    2.80

    Step-by-Step Solution

    Insight: The total service outage percentage is the marginal probability of an outage, found by summing the weighted conditional outage rates.

    Exam route: P(O) = (0.40 0.01) + (0.30 0.05) + (0.30 * 0.03) = 0.028, which is 2.80%.

    Learning route:

    1. We need the overall percentage of time the system is in outage.
    2. This is the marginal probability P(Outage).
    3. Using the Law of Total Probability: P(Outage) = P(Outage|S1)P(S1) + P(Outage|S2)P(S2) + P(Outage|S3)P(S3).
    4. Substitute the values: (0.01 0.40) + (0.05 0.30) + (0.03 * 0.30) = 0.004 + 0.015 + 0.009 = 0.028.
    5. Expressed as a percentage, this is 2.8%. Formatted to 2 decimal places: "2.80".

    Wrong path: Averaging the outage rates (1+5+3)/3 = 3% ignores the fact that the servers are used for different amounts of time.

    Question 9 · Probability Theory · 2022 SUB
    Common Description: Description for the next four questions:
    Dasholytics Inc runs an online analytics dashboard. The company employs three machines—Server 1, Server 2, and Server 3—to serve the data for the dashboard. At any point in time one of these three machines has the job of serving the data, and the other two are kept in standby mode. Dasholytics uses a server scheduler (software) that decides which of the three machines should serve the data at any given point of time. The scheduler switches between data servers without disrupting the data feed to the dashboard.
    The machine which is serving the data sometimes fails to do so; in this case an alert is sent out to the Dasholytics team and they investigate and fix the problem. The time for which a machine fails to serve data is accounted as service outage caused by that machine. For technical reasons the server scheduler is deactivated when there is such an outage: the outage gets over only when the issue with the server is fixed and it is put back online.
    The figures below describe the server usage and outage statistics as compiled over the last one year (365 days). Please use this information to answer the questions that follow.
    Server 1 Server 2 Server 3 0 1 2 3 4 5 Outage Percentage (%) Fig. (a) Server 1: 40% Server 3: 30% Server 2: 30% Fig. (b)
    Figure 1: Figure (a) shows the total service outage caused by each server, as a percentage of the total time that it served data. Figure (b) shows the total time that each server served data, as a percentage of the total duration over which this information was collected. What is the expected number of hours of service outage in an year (365 days)?
    Correct Answer:

    245.28

    Step-by-Step Solution

    Insight: The total outage time is the sum of the outages of each server, which requires weighting each server's conditional outage rate by its marginal usage time.

    Exam route: Calculate P(Outage) = (0.40 0.01) + (0.30 0.05) + (0.30 0.03) = 0.028. Multiply by total hours in a year (365 24 = 8760) to get 245.28.

    Learning route:

    1. Identify the marginal probabilities of each server being active: P(S1) = 0.40, P(S2) = 0.30, P(S3) = 0.30.
    2. Identify the conditional outage probabilities: P(O|S1) = 0.01, P(O|S2) = 0.05, P(O|S3) = 0.03.
    3. Use the Law of Total Probability to find the overall outage probability: P(O) = P(O|S1)P(S1) + P(O|S2)P(S2) + P(O|S3)P(S3) = 0.004 + 0.015 + 0.009 = 0.028.
    4. Convert this probability to expected hours: 0.028 * 8760 hours = 245.28 hours.

    Wrong path: Adding the percentages directly (1% + 5% + 3% = 9%) and applying to 8760 gives 788.4, which ignores the different base serving times.

    Question 10 · Programming · 2022 MSQ
    Consider the following code, in which A is an array indexed from 0 and n is the number of elements in A. function foo(A,n) { L = 0; R = n - 1; while (L <= R) { i = ceil((L + R)/2); if (A[i] < i) { L = i + 1; } else { if (A[i] > i) { R = i - 1; } else { return(i); } } } return(-1); } Here, ceil(x) returns the smallest integer bigger than or equal to the number x.
    If , what will foo(A, 10) return?
    1. A.

      -5

    2. B.

      -1

    3. C.

      4

    4. D.

      6

    Correct Answer:

    ["D"]

    Step-by-Step Solution

    Key idea: This is a Fixed-Point Search using Binary Search. We are looking for an index i such that A[i] = i.

    The array is A = [-5, -4, -3, -2, -1, 4, 6, 8, 10, 12].

    Indices: 0, 1, 2, 3, 4, 5, 6, 7, 8, 9.

    Step 1: Initialization

    L = 0, R = 9.

    Step 2: Iteration 1

    i = ceil((0 + 9) / 2) = ceil(4.5) = 5.

    Check A[5]. A[5] = 4.

    Compare A[5] with i = 5.

    4 < 5.

    Condition A[i] < i is True.

    Update L = i + 1 = 6.

    Current Range: L = 6, R = 9.

    Step 3: Iteration 2

    i = ceil((6 + 9) / 2) = ceil(7.5) = 8.

    Check A[8]. A[8] = 10.

    Compare A[8] with i = 8.

    10 > 8.

    Condition A[i] < i is False.

    Check A[i] > i. True.

    Update R = i - 1 = 7.

    Current Range: L = 6, R = 7.

    Step 4: Iteration 3

    i = ceil((6 + 7) / 2) = ceil(6.5) = 7.

    Check A[7]. A[7] = 8.

    Compare A[7] with i = 7.

    8 > 7.

    Condition A[i] < i is False.

    Check A[i] > i. True.

    Update R = i - 1 = 6.

    Current Range: L = 6, R = 6.

    Step 5: Iteration 4

    L <= R (6 <= 6) is True.

    i = ceil((6 + 6) / 2) = 6.

    Check A[6]. A[6] = 6.

    Compare A[6] with i = 6.

    6 = 6.

    Condition A[i] < i is False.

    Check A[i] > i is False.

    Else block: Return i = 6.

    The function returns the index 6.

    Option D corresponds to 6.

    Other CMI Data Science papers