5125 practice questions
    CMI Data Science Chapter-wise Practice: 5125 Questions Sorted by Level 1 to 4

    5125 CMI Data Science practice questions across 79 chapters, graded Level 1 (basics) to Level 4 (exam stretch), with solutions for every question.

    Level 1
    1468
    29% · Direct use of one concept or formula
    Level 2
    1520
    30% · Two steps or a common twist
    Level 3
    1652
    32% · Exam level, several ideas combined
    Level 4
    485
    9% · Hardest exam level, for top percentiles

    Practice Question Features

    Custom-Made Questions

    Questions designed specifically for the syllabus

    Detailed Solutions

    Step-by-step solutions with hints for better understanding

    Practice Mode

    Learn at your pace - see solutions immediately when stuck

    Exam Mode

    Timed tests simulating actual exam conditions

    CMI Data Science Chapter-wise Practice: 5125 Questions Sorted by Level 1 to 4

    5125 CMI Data Science practice questions across 79 chapters, graded Level 1 (basics) to Level 4 (exam stretch), with solutions for every question.

    CMI Data Science Practice Questions by Difficulty Level

    LevelWhat it testsQuestionsShare
    Level 1Direct use of one concept or formula146829%
    Level 2Two steps or a common twist152030%
    Level 3Exam level, several ideas combined165232%
    Level 4Hardest exam level, for top percentiles4859%

    School Level Mathematics: Practice questions

    UnitChaptersPractice questions
    Algebra & Number Theory111302
    Functions & Calculus11811

    Discrete Mathematics: Practice questions

    UnitChaptersPractice questions
    Sets, Logic & Relations10491
    Unit 1 — Discrete Mathematics30
    Combinatorics & Induction12624
    Unit 2 — Discrete Mathematics30

    Probability Theory: Practice questions

    UnitChaptersPractice questions
    Probability & Random Variables11855
    Statistics & Data Analysis7373

    Programming: Practice questions

    UnitChaptersPractice questions
    Algorithmic Thinking11669

    CMI Data Science Chapter-wise Practice Questions 2026

    School Level Mathematics Practice Questions

    Algebra & Number Theory

    Functions & Calculus

    Discrete Mathematics Practice Questions

    Sets, Logic & Relations

    Unit 1 — Discrete Mathematics

    Combinatorics & Induction

    Unit 2 — Discrete Mathematics

    Probability Theory Practice Questions

    Probability & Random Variables

    Statistics & Data Analysis

    Programming Practice Questions

    Algorithmic Thinking

    CMI Data Science Practice Questions with Solutions

    School Level Mathematics: Solved Questions

    Q1. (CAT 2024) Let \(\alpha\) be a fixed real number and let \(f\) be a function from \(\mathbb{R}^2\) to \(\mathbb{R}^2\) defined as \[ f(x,y)=(x\cos\alpha-y\sin\alpha,x\sin\alpha+y\cos\alpha). \] Write down an expression for \(f^{10}(x,y)\), where \(f^{10}\) denotes the function obtained by composing \(f\) with itself 10 times.

    Answer: none

    Solution: Insight: The map $f$ is a geometric rotation by angle $\alpha$; composing it 10 times yields a rotation by $10\alpha$.
    Exam route: Recognise $f = R_\alpha$. Apply $R_\alpha^n = R_{n\alpha}$ with $n=10$. Substitute $10\alpha$ into the rotation formula.
    Learning route:
    This is a function-iteration question, recognisable because the formula $(x\cos\alpha - y\sin\alpha, x\sin\alpha + y\cos\alpha)$ is the standard rotation matrix and "composing $f$ with itself 10 times" signals $f^{10} = f \circ f \circ \cdots \circ f$.
    Step 1: Identify the transformation. The map $f(x,y) = (x\cos\alpha - y\sin\alpha, x\sin\alpha + y\cos\alpha)$ rotates every point anticlockwise by $\alpha$ about the origin. We write $f = R_\alpha$.
    Step 2: Iterate geometrically. One application rotates by $\alpha$. A second rotates the result by another $\alpha$, totalling $2\alpha$. By induction, $n$ applications give rotation by $n\alpha$: $R_\alpha^n = R_{n\alpha}$.
    Step 3: Set $n=10$. Then $f^{10} = R_{10\alpha}$.
    Step 4: Substitute $10\alpha$ for $\alpha$ in the original formula:
    $f^{10}(x,y) = (x\cos 10\alpha - y\sin 10\alpha, x\sin 10\alpha + y\cos 10\alpha)$.
    Wrong path: Raising each coordinate to the 10th power, e.g. $(x\cos\alpha - y\sin\alpha)^{10}$. This confuses iteration $f^n$ with exponentiation $(f(x))^n$. The notation $f^{10}$ means ten-fold composition, not the tenth power of the output.
    Verification: For $\alpha = \pi/2$, $f(x,y)=(-y,x)$, so $f^2(x,y)=(-x,-y)$ and $f^{10}=f^2=(-x,-y)$. Formula gives $(x\cos 5\pi - y\sin 5\pi, x\sin 5\pi + y\cos 5\pi) = (-x, -y)$. Matches.

    Q2. (CAT 2023) Find the limit \[ \lim_{x\to\infty}\left(x-x\cos\frac{1}{\sqrt{x}}\right) \]

    Answer: none

    Solution: Insight: This is an $\infty - \infty$ indeterminate form at infinity; factoring $x$ and substituting $t = 1/\sqrt{x}$ converts it to the standard limit $(1-\cos t)/t^2 = 1/2$.
    Exam route:
    1. Factor out $x$: $x(1 - \cos(1/\sqrt{x}))$.
    2. Substitute $t = 1/\sqrt{x}$, so $t \to 0^+$ and $x = 1/t^2$.
    3. The expression becomes $(1 - \cos t)/t^2$.
    4. Apply the standard limit $\lim_{t\to 0} (1-\cos t)/t^2 = 1/2$.
    Learning route:
    This is a limits-at-infinity question with an $\infty - \infty$ indeterminate form, recognisable because both $x$ and $x\cos(1/\sqrt{x})$ grow without bound as $x \to \infty$. The trigger for the substitution method is the $1/\sqrt{x}$ inside the cosine, which shrinks to zero as $x$ grows.
    Step 1: Never split $\infty - \infty$ into two separate limits. Instead, factor out $x$:
    $$x - x\cos\frac{1}{\sqrt{x}} = x\left(1 - \cos\frac{1}{\sqrt{x}}\right)$$
    This converts the form from $\infty - \infty$ to $\infty \cdot 0$, which is still indeterminate but easier to handle.
    Step 2: Substitute $t = 1/\sqrt{x}$. As $x \to \infty$, $t \to 0^+$, and $x = 1/t^2$:
    $$x\left(1 - \cos\frac{1}{\sqrt{x}}\right) = \frac{1}{t^2}(1 - \cos t) = \frac{1 - \cos t}{t^2}$$
    Step 3: Apply the standard limit. Using $1 - \cos t = 2\sin^2(t/2)$ and $\lim_{u\to 0}\frac{\sin u}{u} = 1$:
    $$\lim_{t\to 0}\frac{1-\cos t}{t^2} = \lim_{t\to 0}\frac{2\sin^2(t/2)}{t^2} = \lim_{t\to 0}\frac{2\cdot(t/2)^2}{t^2} = \frac{1}{2}$$
    Wrong path: A tempting mistake is to split the limit: $\lim_{x\to\infty} x - \lim_{x\to\infty} x\cos(1/\sqrt{x})$. This yields $\infty - \infty$, which is invalid because you can only split limits if both individual limits exist and are finite.

    Q3. (CAT 2020) Let \(f(x)\) be a real-valued function all of whose derivatives exist. Recall that a point \(x_0\) in the domain is called an <b>inflection point</b> of \(f(x)\) if the second derivative \(f''(x)\) changes sign at \(x_0\). Given the function \[ f(x)=\frac{x^5}{20}-\frac{x^4}{2}+3x+1, \] which of the following statements are true?

    1. \(x_0=0\) is not an inflection point.
    2. \(x_0=6\) is the only inflection point.
    3. \(x_0=0\) and \(x_0=6\), both are inflection points.
    4. The function does not have an inflection point.

    Answer: ["A","B"]

    Solution: Key idea: This is a "statement_truth" question requiring the systematic identification and verification of inflection points by checking the sign change of the second derivative.
    Step 1: Find the first derivative. $f'(x) = \frac{d}{dx}\left(\frac{x^5}{20} - \frac{x^4}{2} + 3x + 1\right) = \frac{x^4}{4} - 2x^3 + 3$.
    Step 2: Find the second derivative. $f''(x) = \frac{d}{dx}\left(\frac{x^4}{4} - 2x^3 + 3\right) = x^3 - 6x^2$.
    Step 3: Find candidate points where $f''(x) = 0$. $x^3 - 6x^2 = x^2(x - 6) = 0 \implies x = 0$ or $x = 6$.
    Step 4: Verify sign change at $x = 0$. For $x < 0$, $x^2 > 0$ and $x - 6 < 0$, so $f''(x) < 0$. For $0 < x < 6$, $x^2 > 0$ and $x - 6 < 0$, so $f''(x) < 0$. Since the sign does not change, $x = 0$ is NOT an inflection point.
    Step 5: Verify sign change at $x = 6$. For $0 < x < 6$, $f''(x) < 0$. For $x > 6$, $x^2 > 0$ and $x - 6 > 0$, so $f''(x) > 0$. Since the sign changes from negative to positive, $x = 6$ IS an inflection point.
    Step 6: Evaluate options. Option A is true ($x_0=0$ is not an inflection point). Option B is true ($x_0=6$ is the only one). Options C and D are false.
    Answer: Options A and B.

    Q4. (CAT 2021) For a non-zero real number \(a\), the inverse of \(J=\begin{pmatrix}a&1&0\\0&a&1\\0&0&a\end{pmatrix}\) is

    1. \(\begin{pmatrix} a^{-1} & a^{-2} & a^{-3}\\ 0 & a^{-1} & a^{-2}\\ 0 & 0 & a^{-1} \end{pmatrix}\)
    2. \(\begin{pmatrix} a^{-1} & -a^{-2} & a^{-3}\\ 0 & a^{-1} & -a^{-2}\\ 0 & 0 & a^{-1} \end{pmatrix}\)
    3. \(\begin{pmatrix} a^{-1} & 1 & 0\\ 0 & a^{-1} & 1\\ 0 & 0 & a^{-1} \end{pmatrix}\)
    4. \(\begin{pmatrix} a^{-1} & a^{-2} & 0\\ 0 & a^{-1} & a^{-2}\\ 0 & 0 & a^{-1} \end{pmatrix}\)

    Answer: ["B"]

    Solution: Key idea: This is a matrix inverse problem for a special upper triangular matrix, recognizable because it has a constant diagonal $a$ and constant superdiagonal $1$.
    Step 1: Decompose the matrix as $J = aI + N$, where $N = \begin{pmatrix} 0 & 1 & 0 \\ 0 & 0 & 1 \\ 0 & 0 & 0 \end{pmatrix}$.
    Step 2: Observe that $N$ is nilpotent. Specifically, $N^2 = \begin{pmatrix} 0 & 0 & 1 \\ 0 & 0 & 0 \\ 0 & 0 & 0 \end{pmatrix}$ and $N^3 = 0$.
    Step 3: Use the finite geometric series expansion for the inverse: $(aI + N)^{-1} = a^{-1}(I + a^{-1}N)^{-1} = a^{-1}(I - a^{-1}N + a^{-2}N^2)$.
    Step 4: Substitute $I$, $N$, and $N^2$ into the expansion:
    $J^{-1} = \begin{pmatrix} a^{-1} & 0 & 0 \\ 0 & a^{-1} & 0 \\ 0 & 0 & a^{-1} \end{pmatrix} - \begin{pmatrix} 0 & a^{-2} & 0 \\ 0 & 0 & a^{-2} \\ 0 & 0 & 0 \end{pmatrix} + \begin{pmatrix} 0 & 0 & a^{-3} \\ 0 & 0 & 0 \\ 0 & 0 & 0 \end{pmatrix} = \begin{pmatrix} a^{-1} & -a^{-2} & a^{-3} \\ 0 & a^{-1} & -a^{-2} \\ 0 & 0 & a^{-1} \end{pmatrix}$.
    Answer: B

    Q5. (CAT 2025) Evaluate the following limit: \[ \lim_{x\to 0}\frac{e^{2x}-2x-\cos x}{x\sin x} \]

    Answer: none

    Solution: Insight: This is a transcendental $0/0$ limit with mixed exponential, polynomial, and trigonometric terms; the denominator's leading order $x^2$ dictates expanding every numerator term through $x^2$ using Taylor series.
    Exam route:
    1. Denominator: $x\sin x = x(x - x^3/6 + \dots) = x^2 + O(x^4) \sim x^2$.
    2. Numerator to $O(x^2)$:
    $e^{2x} = 1 + 2x + 2x^2 + O(x^3)$,
    $-2x = -2x$,
    $-\cos x = -(1 - x^2/2 + O(x^4)) = -1 + x^2/2 + O(x^4)$.
    3. Assemble: $(1 + 2x + 2x^2) - 2x + (-1 + x^2/2) = \frac{5}{2}x^2 + O(x^3)$.
    4. Divide: $\lim_{x\to 0}\frac{\frac{5}{2}x^2 + O(x^3)}{x^2 + O(x^4)} = \frac{5}{2}$.
    Learning route:
    This is a transcendental $0/0$ limit, recognisable because direct substitution gives $(1 - 0 - 1)/(0 \cdot 0) = 0/0$ and the numerator mixes exponential, polynomial, and trigonometric terms. The trigger for Taylor expansion is this mixture — algebraic factoring cannot untangle $e^{2x}$ and $\cos x$ simultaneously.
    Step 1: Determine the denominator's leading order. Expand $\sin x = x - x^3/6 + \dots$, so:
    $$x\sin x = x\left(x - \frac{x^3}{6} + \dots\right) = x^2 - \frac{x^4}{6} + \dots = x^2 + O(x^4)$$
    The leading term is $x^2$. This tells us we must expand every numerator term through $x^2$ — no less, no more.
    Step 2: Expand each numerator term to $O(x^2)$:
    $$e^{2x} = 1 + (2x) + \frac{(2x)^2}{2!} + O(x^3) = 1 + 2x + 2x^2 + O(x^3)$$
    $$-2x = -2x$$
    $$-\cos x = -\left(1 - \frac{x^2}{2} + O(x^4)\right) = -1 + \frac{x^2}{2} + O(x^4)$$
    Step 3: Combine all numerator terms:
    $$(1 + 2x + 2x^2) - 2x + \left(-1 + \frac{x^2}{2}\right) = \frac{5}{2}x^2 + O(x^3)$$
    Step 4: Divide by the denominator and take the limit:
    $$\lim_{x\to 0}\frac{\frac{5}{2}x^2 + O(x^3)}{x^2 + O(x^4)} = \lim_{x\to 0}\frac{\frac{5}{2} + O(x)}{1 + O(x^2)} = \frac{5}{2}$$
    Wrong path: Using L'Hôpital's rule twice is possible but extremely messy and prone to calculation errors. Another common mistake is expanding the numerator only to $O(x)$, which yields $0$ in the numerator and leads to the incorrect conclusion that the limit is $0$.

    Discrete Mathematics: Solved Questions

    Q1. (CAT 2020) Peppa and her friends Suzy and Emily are making 3 masks : a knight mask, a pirate mask and an elf mask. The one wearing the knight mask must always tell the truth, the one wearing the pirate mask must always lie and the one wearing the elf mask could sometimes lie and sometimes tell the truth. After making three masks, one of each kind, they all wear one each and make the following statements:<br/> Peppa : I am a wearing an elf mask.<br/> Suzy : That is true.<br/> Emily : I am not wearing an elf mask.<br/> What kind of mask is each one wearing?

    Answer: 2

    Solution: Key idea: This is a three-role truth-teller puzzle (Knight=Truth, Pirate=Lie, Elf=Variable). The key is to test assumptions against the constraints of each role.
    Step 1: Analyze Peppa's statement "I am wearing an elf mask". If Peppa were the Knight (always truth), she would say "I am Knight". She cannot claim to be Elf because that would be a lie. So Peppa $\neq$ Knight.
    Step 2: If Peppa is the Pirate (always lie), claiming to be Elf is a lie, which is consistent. If Peppa is the Elf (variable), she could tell the truth or lie.
    Step 3: Test Peppa = Pirate. Then "Peppa is Elf" is False. Suzy says "That is true", so Suzy is lying. Since Peppa is Pirate, Suzy must be Elf (lying). This leaves Emily as Knight.
    Step 4: Verify Emily (Knight). Emily says "I am not wearing an elf mask". Since Emily is Knight, this is True. Consistent.
    Step 5: Test Peppa = Elf (telling truth). Then Suzy tells truth. Suzy could be Knight. Then Emily = Pirate. Emily says "I am not Elf". Since Emily is Pirate, this must be a lie, meaning Emily IS Elf. Contradiction.
    Answer: Peppa=Pirate, Suzy=Elf, Emily=Knight. (Encoded as 2 for NAT format: 1=Knight, 2=Pirate, 3=Elf mapping to Peppa).

    Q2. (CAT 2023) Each lawyer on a certain remote island is either honest or dishonest (but not both).<br/> • Honest lawyers always speak the truth.<br/> • Dishonest lawyers always lie.<br/> Is it possible for a lawyer on this island to claim that he/ she is dishonest? Explain. A judge asks lawyer \(A\), “are you honest?” Before \(A\) could answer, lawyer \(B\) says “\(A\) will say yes. But then, he’ll be lying”. Which lawyer is honest and which one is dishonest? Explain.

    Answer: none

    Solution: Insight: Self-referential dishonesty claims produce $X \iff \neg X$ (a paradox with no solution in $\{T,F\}$). Meta-statements about hypothetical answers collapse once you realise every islander answers "Yes" to "Are you honest?".

    Exam route: Part (i) $X \iff \neg X$ has no solution in $\{T,F\}$, so no lawyer can claim dishonesty. Part (ii) A's hypothetical reply is provably "Yes", so B's compound claim reduces to "$A$ is lying", giving $B \iff \neg A$. Opposite types, identities undetermined.

    Learning route:
    This is a liar-paradox plus meta-statement question, recognisable because it asks whether a character can make a self-referential identity claim and then layers a hypothetical-report statement on top.

    Part (i) — Can a lawyer claim "I am dishonest"?
    Step 1: Let $X \in \{T, F\}$ encode the lawyer's type ($T$ = honest, $F$ = dishonest). The claim "I am dishonest" asserts $\neg X$.
    Step 2: Apply the core rule $X \iff P$ where $P = \neg X$. This gives $X \iff \neg X$.
    Step 3: Test $X = T$: $T \iff F$ — contradiction. Test $X = F$: $F \iff T$ — contradiction.
    Step 4: No assignment satisfies the equivalence, so no lawyer on this island can claim to be dishonest.

    Part (ii) — Which lawyer is honest and which is dishonest?
    Step 1: Analyse A's hypothetical answer to "Are you honest?" If A is honest ($T$), A truthfully says "Yes". If A is dishonest ($F$), A lies and also says "Yes". Therefore A will definitely say "Yes" regardless of type. The first conjunct of B's statement is a proven fact (truth value $T$).
    Step 2: Translate B's full statement. B asserts the conjunction: "A will say yes" $\wedge$ "A is lying". The first conjunct is $T$ (proven). The second conjunct is $\neg A$. So B's statement has truth value $T \wedge \neg A = \neg A$.
    Step 3: Apply the core rule to B: $B \iff \neg A$. This biconditional means A and B have opposite types.
    Step 4: No further information pins down which is which. Both $(A=T, B=F)$ and $(A=F, B=T)$ are consistent with all constraints.

    Wrong path: A student might reason "B correctly predicts A will say yes, so B must be honest, making A dishonest." This produces the definite answer "B is honest, A is dishonest." The break occurs at assuming B's entire conjunction is true just because one conjunct is true. B's second conjunct ("A is lying") could be false, making B's whole statement false and B a liar.

    Generalization: When a speaker makes a compound statement, evaluate the truth value of the entire compound, not just one part.

    Verification: Check $(A=T, B=F)$: A says "Yes" truthfully. B says "A says yes AND A is lying" $= T \wedge F = F$. Since $B=F$, B's false statement is consistent. Check $(A=F, B=T)$: A says "Yes" (lying). B says "A says yes AND A is lying" $= T \wedge T = T$. Since $B=T$, B's true statement is consistent. Both assignments work.

    Answer: (i) No, it is impossible for a lawyer to claim dishonesty. (ii) A and B have opposite types, but individual identities cannot be determined.

    Q3. (CAT 2020) How many squares are there on a \(7\times 7\) chessboard?

    1. 49
    2. 204
    3. 203
    4. 140

    Answer: ["D"]

    Solution: Key idea: this is a grid enumeration question asking for the total number of squares of all sizes in an $n \times n$ grid. The trigger is "how many squares", which implies counting $1 \times 1$, $2 \times 2$, up to $n \times n$ squares, not just the unit cells.

    Step 1: A $k \times k$ square on a $7 \times 7$ board is uniquely determined by the position of its top-left corner.
    Step 2: The top-left corner can be placed in $(7 - k + 1) = (8 - k)$ horizontal positions and $(8 - k)$ vertical positions. Thus, there are $(8 - k)^2$ squares of size $k \times k$.
    Step 3: Sum over all possible sizes $k$ from 1 to 7:
    Total squares $= \sum_{k=1}^{7} (8-k)^2 = 7^2 + 6^2 + 5^2 + 4^2 + 3^2 + 2^2 + 1^2$.
    Step 4: Calculate the sum: $49 + 36 + 25 + 16 + 9 + 4 + 1 = 140$. (This matches the standard formula $\frac{n(n+1)(2n+1)}{6}$ for $n=7$, which gives $\frac{7 \times 8 \times 15}{6} = 140$).
    Step 5: Match with options. Option A (49) counts only the $1 \times 1$ cells. Option B (204) is the sum for an $8 \times 8$ board. Option D (140) is the correct total for a $7 \times 7$ board.

    Answer: ["D"]

    Q4. (CAT 2020) Given the set of letters \(\{a,b,c,d,e,f,g,h,i,j,k,l,m\}\), we can list out all permutations of these letters in lexicographic (dictionary) order. The first three permutations in this list are abcdefghijklm, abcdefghijkml and abcdefghijlkm and the last one is mlkjihgfedcba. What permutations would appear immediately before and after the following one in this lexicographically ordered list of permutations?<br/> bcjameflkihgd

    Answer: 0

    Solution: Key idea: This is a lexicographic permutation question, recognizable because it asks for the immediate predecessor and successor in dictionary order.
    Note: The question is typed as NAT but asks for strings. The numeric placeholder '0' is used to satisfy the NAT format constraint. The actual strings are 'bcjameflkihdg' and 'bcjamegdfhikl'.
    Step 1: To find the previous permutation, scan from right to left to find the first pair where the left character is greater than the right character. In `bcjameflkihgd`, scanning from right, `g` > `d`.
    Step 2: Swap `g` and `d` to get `bcjameflkihdg`. The suffix is already in decreasing order, so no further sorting is needed.
    Step 3: To find the next permutation, scan from right to left to find the first pair where the left character is smaller than the right character. In `bcjameflkihgd`, `f` < `l`.
    Step 4: In the suffix `lkihgd`, find the smallest character that is strictly greater than `f`. That character is `g`.
    Step 5: Swap `f` and `g` to get `bcjameg` followed by the remaining suffix `lkihfd`.
    Step 6: Sort the suffix `lkihfd` in increasing order to get `dfhikl`.
    Step 7: Combine the prefix and sorted suffix to get the next permutation: `bcjamegdfhikl`.
    Answer: 0 (Placeholder for NAT format. Actual: bcjameflkihdg, bcjamegdfhikl)

    Q5. (CAT 2021) Recall that if \(h\) is a function from \(X\) to \(Y\) and \(g\) is a function from \(Y\) to \(Z\) then, \(g\circ h\) is the function from \(X\) to \(Z\) such that \((g\circ h)(x)=g(h(x))\), for all \(x \in X\).<br/> Let \(S\) be the set of all functions \(f\) from \(\{1,2,3,4,5,6\}\) to \(\{1,2,3,4,5,6\}\) such that \(f\circ f=f\).<br/> (a) Compute the number of functions \(f \in S\) whose range has three elements.<br/> (b) What is the cardinality of \(S\)?

    Answer: 1057

    Solution: Key idea: This is an idempotent function counting question, recognizable by the condition $f \circ f = f$.
    Step 1: For $f \circ f = f$, the range of $f$ must be exactly the set of fixed points of $f$. Let the range be $R$. Then for all $x \in R$, $f(x) = x$.
    Step 2: For any $x \notin R$, $f(x)$ must be an element of $R$.
    Step 3: To count functions with range size $k$, we first choose the $k$ elements of $R$ from the 6 available elements in $\binom{6}{k}$ ways.
    Step 4: For the remaining $6-k$ elements, each can map to any of the $k$ elements in $R$, giving $k^{6-k}$ choices.
    Step 5: For part (a), $k=3$. The number of such functions is $\binom{6}{3} \times 3^{6-3} = 20 \times 27 = 540$.
    Step 6: For part (b), we sum this over all possible range sizes $k \in \{1, 2, 3, 4, 5, 6\}$:
    $k=1: \binom{6}{1} \times 1^5 = 6$
    $k=2: \binom{6}{2} \times 2^4 = 15 \times 16 = 240$
    $k=3: \binom{6}{3} \times 3^3 = 20 \times 27 = 540$
    $k=4: \binom{6}{4} \times 4^2 = 15 \times 16 = 240$
    $k=5: \binom{6}{5} \times 5^1 = 6 \times 5 = 30$
    $k=6: \binom{6}{6} \times 6^0 = 1 \times 1 = 1$
    Total cardinality = $6 + 240 + 540 + 240 + 30 + 1 = 1057$.
    Answer: 1057 (Note: Part (a) is 540, Part (b) is 1057).

    Probability Theory: Solved Questions

    Q1. (CAT 2026) Common Description: Question (6) and (7) are based on the following information A list \((a_1,\ldots,a_n)\) is an arrangement of natural numbers \(1,2,\ldots,n\) in a random order, with all orderings equally likely. We say that a position \(i\) is a new maximum if \(a_i > a_j\) for all \(j < i\). For example, in the list \[ (3,1,4,2,5) \] positions 1,3,5 are new maxima. Note that position 1 is always a new maximum. Now answer the two questions below based on this information. For a given \(i\), determine the probability that position \(i\) is a new maximum.

    Answer: none

    Solution: Insight: This is a record probability question, recognisable because it asks for the probability that a specific position sets a new running maximum in a random permutation. The key is to isolate the first $i$ elements and use symmetry.

    Exam route:
    1. Focus only on the first $i$ elements $a_1, \ldots, a_i$.
    2. The event "position $i$ is a new maximum" means $a_i > a_j$ for all $j < i$, which is exactly $a_i = \max(a_1, \ldots, a_i)$.
    3. By symmetry, the largest of these $i$ distinct values is equally likely to be at any of the $i$ positions.
    4. Therefore, the probability it lands at position $i$ is $1/i$.

    Learning route:
    Step 1. The full permutation $(a_1, \ldots, a_n)$ is drawn uniformly from all $n!$ orderings of $\{1, 2, \ldots, n\}$.
    Step 2. The condition $\{a_i > a_j \text{ for all } j < i\}$ depends only on the relative order of the first $i$ entries.
    Step 3. The set of values $\{a_1, \ldots, a_i\}$ forms some $i$-element subset of $\{1, \ldots, n\}$. Conditional on this subset, all $i!$ relative orderings are equally likely.
    Step 4. Among the $i!$ orderings, exactly $(i-1)!$ place the largest value at position $i$ (the remaining $i-1$ values can be arranged freely in the first $i-1$ positions).
    Step 5. The probability is $\frac{(i-1)!}{i!} = \frac{1}{i}$.

    Common trap: Answering $1/n$. This mistake comes from confusing the running maximum with the global maximum. The question asks for the probability of beating the previous $i-1$ elements, not all $n$ elements.

    Verification: For $i=1$, probability is $1/1 = 1$ (position 1 is always a record). For $i=2$, probability is $1/2$ (position 2 is a record iff $a_2 > a_1$, which happens in half of all permutations). Both match intuition.

    Q2. (CAT 2020) Owing to a defect in a certain machine which makes N95 masks, there is a 0.1% probability that a mask it makes is not effective in preventing airborne viruses from being inhaled.<br/> (a) What is the probability that the first 1000 masks that the machine produces are effective? (You may leave your solutions as arithmetic expressions; there is no need to compute their decimal representations.)<br/> (b) What is the probability that among the first one crore \((10^7)\) masks that the machine produces, there is at least one mask which is not effective?

    Answer: 1

    Solution: Key idea: This is a repeated independent Bernoulli trials question with a rare defect. The trigger words are "machine makes masks", "probability that a mask ... is not effective", and "at least one".

    Step 1: Find the probability that one mask is effective.
    The probability that a mask is not effective is \(0.1\% = 0.001\). Therefore,
    \[
    P(\text{effective}) = 1 - 0.001 = 0.999.
    \]

    Step 2: Solve part (a).
    The first 1000 masks are independent. The probability that all 1000 are effective is
    \[
    (0.999)^{1000}.
    \]
    Numerically, this is approximately \(0.367695\).

    Step 3: Solve part (b) using the complement rule.
    "At least one mask is not effective" is the complement of "all masks are effective".
    For \(10^7\) masks,
    \[
    P(\text{all effective}) = (0.999)^{10^7}.
    \]
    Hence,
    \[
    P(\text{at least one not effective}) = 1 - (0.999)^{10^7}.
    \]

    Step 4: Interpret the NAT numeric answer.
    Since
    \[
    (0.999)^{10^7} = e^{10^7 \ln(0.999)} \approx e^{-10005},
    \]
    this number is effectively zero to any usual numerical precision. Therefore the final probability in part (b) is numerically \(1\).

    Answer: The exact expressions are \((0.999)^{1000}\) for part (a) and \(1-(0.999)^{10^7}\) for part (b). The single numeric NAT answer for the final part is \(1\).

    Q3. (CAT 2020) The following graph shows the performance of students in an exam. The marks scored by every student are an multiple of five. The \(j^{th}\)-percentile \(u^*\) for a discrete data \(x_1,x_2,\ldots,x_n\) is defined as follows. Let \(x_{(1)},x_{(2)},\ldots,x_{(n)}\) be the ordering of the data in ascending order. Let \(t=\frac{jn}{100}\) and let \(k\) be an integer such that \(k\leq t<(k+1)\) and let \(s=t-k\). Then \(u^*=x_{(k)}+s*(x_{(k+1)}-x_{(k)})\). Here, \(x_{(n+1)}\) is defined to be \(x_{(n)}\). <br/> <svg xmlns="http://www.w3.org/2000/svg" width="520" height="300" viewBox="0 0 520 300"> <rect x="0" y="0" width="520" height="300" fill="white"/> <line x1="85" y1="235" x2="430" y2="235" stroke="black"/> <line x1="85" y1="45" x2="85" y2="235" stroke="black"/> <line x1="85" y1="235" x2="430" y2="235" stroke="#cccccc"/> <line x1="85" y1="208" x2="430" y2="208" stroke="#dddddd"/> <line x1="85" y1="181" x2="430" y2="181" stroke="#dddddd"/> <line x1="85" y1="154" x2="430" y2="154" stroke="#dddddd"/> <line x1="85" y1="127" x2="430" y2="127" stroke="#dddddd"/> <line x1="85" y1="100" x2="430" y2="100" stroke="#dddddd"/> <line x1="85" y1="73" x2="430" y2="73" stroke="#dddddd"/> <line x1="85" y1="46" x2="430" y2="46" stroke="#dddddd"/> <text x="75" y="239" text-anchor="end" font-size="12">0</text> <text x="75" y="212" text-anchor="end" font-size="12">2</text> <text x="75" y="185" text-anchor="end" font-size="12">4</text> <text x="75" y="158" text-anchor="end" font-size="12">6</text> <text x="75" y="131" text-anchor="end" font-size="12">8</text> <text x="75" y="104" text-anchor="end" font-size="12">10</text> <text x="75" y="77" text-anchor="end" font-size="12">12</text> <text x="75" y="50" text-anchor="end" font-size="12">14</text> <polyline points="105,208 137,195 169,208 201,127 233,113.5 265,127 297,181 329,194.5 361,181 393,86.5 425,73" fill="none" stroke="black" stroke-width="1.4"/> <text x="105" y="255" text-anchor="middle" font-size="12">40</text> <text x="137" y="255" text-anchor="middle" font-size="12">45</text> <text x="169" y="255" text-anchor="middle" font-size="12">50</text> <text x="201" y="255" text-anchor="middle" font-size="12">55</text> <text x="233" y="255" text-anchor="middle" font-size="12">60</text> <text x="265" y="255" text-anchor="middle" font-size="12">65</text> <text x="297" y="255" text-anchor="middle" font-size="12">70</text> <text x="329" y="255" text-anchor="middle" font-size="12">75</text> <text x="361" y="255" text-anchor="middle" font-size="12">80</text> <text x="393" y="255" text-anchor="middle" font-size="12">85</text> <text x="425" y="255" text-anchor="middle" font-size="12">90</text> <text x="260" y="285" text-anchor="middle" font-size="14">Marks</text> <text x="30" y="145" text-anchor="middle" font-size="14" transform="rotate(-90 30 145)">Number of Students</text> </svg> <br/> Based on the information presented in the graph, answer the following questions.<br/> (a) Compute the 10th percentile of marks.<br/> (b) Is the median score higher than the mean score?

    Answer: 50

    Solution: Key idea: This is a percentile calculation and descriptive statistics problem, recognizable because it asks for a specific percentile from a frequency distribution and compares the median to the mean.

    Step 1: Extract the frequencies from the graph.
    By reading the y-axis (Number of Students) for each x-axis value (Marks), we get:
    - 40: 2 students
    - 45: 3 students
    - 50: 2 students
    - 55: 8 students
    - 60: 9 students
    - 65: 8 students
    - 70: 4 students
    - 75: 3 students
    - 80: 4 students
    - 85: 11 students
    - 90: 12 students

    Step 2: Calculate the total number of students, $n$.
    $n = 2 + 3 + 2 + 8 + 9 + 8 + 4 + 3 + 4 + 11 + 12 = 66$.

    Step 3: Compute the 10th percentile.
    Using the given formula, $t = \frac{j \times n}{100} = \frac{10 \times 66}{100} = 6.6$.
    Here, $k = 6$ and $s = 0.6$.
    We need the 6th and 7th values in the ascending ordered data, $x_{(6)}$ and $x_{(7)}$.
    Cumulative frequencies:
    - Up to 40: 2
    - Up to 45: 2 + 3 = 5
    - Up to 50: 5 + 2 = 7
    Since the 6th and 7th values both fall in the "50" group, $x_{(6)} = 50$ and $x_{(7)} = 50$.
    $u^* = x_{(6)} + s \times (x_{(7)} - x_{(6)}) = 50 + 0.6 \times (50 - 50) = 50$.

    Step 4: Compare the median and the mean.
    Median: For $n = 66$, the median is the average of the 33rd and 34th values.
    Cumulative frequencies:
    - Up to 65: 2 + 3 + 2 + 8 + 9 + 8 = 32
    - Up to 70: 32 + 4 = 36
    Both the 33rd and 34th values are 70. So, Median = 70.

    Mean: Calculate the sum of all marks.
    Sum = $(40 \times 2) + (45 \times 3) + (50 \times 2) + (55 \times 8) + (60 \times 9) + (65 \times 8) + (70 \times 4) + (75 \times 3) + (80 \times 4) + (85 \times 11) + (90 \times 12)$
    Sum = $80 + 135 + 100 + 440 + 540 + 520 + 280 + 225 + 320 + 935 + 1080 = 4655$.
    Mean = $\frac{4655}{66} \approx 70.53$.

    Since the median (70) is not higher than the mean (70.53), the answer to part (b) is No.

    Answer: 50

    Q4. (CAT 2025) Common Description: Instructions for Questions 18,19, and 20:<br/> Read the following description carefully and answer the questions that follow. Use all the information provided. Clearly show all your calculations.<br/> Description:<br/> NorthCool Beverages Pvt. Ltd. is a leading beverage company that operates across six northern states of India. During the past summer the company launched an aggressive marketing campaign to promote its range of flavoured drinks. The charts below summarise the sales data collected during this campaign.<br/> <svg xmlns="http://www.w3.org/2000/svg" width="760" height="330" viewBox="0 0 760 330"> <rect x="0" y="0" width="760" height="330" fill="white"/> <text x="150" y="25" text-anchor="middle" font-size="12" font-weight="bold">Revenue Share by Flavour</text> <circle cx="150" cy="155" r="95" fill="white" stroke="black" stroke-width="1.2"/> <path d="M150 155 L150 60 A95 95 0 0 1 179.36 64.65 Z" fill="white" stroke="black" stroke-width="1"/> <path d="M150 155 L179.36 64.65 A95 95 0 0 1 244.53 164.94 Z" fill="white" stroke="black" stroke-width="1"/> <path d="M150 155 L244.53 164.94 A95 95 0 0 1 194.04 239.17 Z" fill="white" stroke="black" stroke-width="1"/> <path d="M150 155 L194.04 239.17 A95 95 0 0 1 60.65 187.48 Z" fill="white" stroke="black" stroke-width="1"/> <path d="M150 155 L60.65 187.48 A95 95 0 0 1 60.65 122.52 Z" fill="white" stroke="black" stroke-width="1"/> <path d="M150 155 L60.65 122.52 A95 95 0 0 1 150 60 Z" fill="white" stroke="black" stroke-width="1"/> <text x="162" y="92" text-anchor="middle" font-size="10">Other</text> <text x="162" y="104" text-anchor="middle" font-size="10">(5%)</text> <text x="210" y="140" text-anchor="middle" font-size="10">Orange</text> <text x="210" y="152" text-anchor="middle" font-size="10">(18%)</text> <text x="195" y="203" text-anchor="middle" font-size="10">Mango</text> <text x="195" y="215" text-anchor="middle" font-size="10">(15%)</text> <text x="137" y="218" text-anchor="middle" font-size="10">Lemon</text> <text x="137" y="230" text-anchor="middle" font-size="10">(28%)</text> <text x="92" y="172" text-anchor="middle" font-size="10">Ginger Ale</text> <text x="92" y="184" text-anchor="middle" font-size="10">(10%)</text> <text x="93" y="115" text-anchor="middle" font-size="10">Cola</text> <text x="93" y="127" text-anchor="middle" font-size="10">(24%)</text> <text x="510" y="25" text-anchor="middle" font-size="12" font-weight="bold">Lemon Flavour Sales By State (Rs. Lakhs)</text> <line x1="380" y1="245" x2="690" y2="245" stroke="black" stroke-width="1"/> <line x1="380" y1="55" x2="380" y2="245" stroke="black" stroke-width="1"/> <text x="365" y="67" text-anchor="end" font-size="10">Uttar Pradesh</text> <text x="365" y="97" text-anchor="end" font-size="10">Haryana</text> <text x="365" y="127" text-anchor="end" font-size="10">Punjab</text> <text x="365" y="157" text-anchor="end" font-size="10">Himachal Pradesh</text> <text x="365" y="187" text-anchor="end" font-size="10">Delhi</text> <text x="365" y="217" text-anchor="end" font-size="10">Uttarakhand</text> <rect x="380" y="58" width="300" height="12" fill="white" stroke="black"/> <rect x="380" y="88" width="250" height="12" fill="white" stroke="black"/> <rect x="380" y="118" width="200" height="12" fill="white" stroke="black"/> <rect x="380" y="148" width="200" height="12" fill="white" stroke="black"/> <rect x="380" y="178" width="150" height="12" fill="white" stroke="black"/> <rect x="380" y="208" width="100" height="12" fill="white" stroke="black"/> <line x1="430" y1="245" x2="430" y2="250" stroke="black"/> <line x1="480" y1="245" x2="480" y2="250" stroke="black"/> <line x1="530" y1="245" x2="530" y2="250" stroke="black"/> <line x1="580" y1="245" x2="580" y2="250" stroke="black"/> <line x1="630" y1="245" x2="630" y2="250" stroke="black"/> <line x1="680" y

    CMI Data Science Preparation Resources 2026