fix for add from pool bug (long loading time) mantis 47724 - #11821
fix for add from pool bug (long loading time) mantis 47724#11821mglaubitz wants to merge 6 commits into
Conversation
… index
The 'Add from pool' question browser (ilObjTestGUI -> ilTestQuestionBrowserTableGUI
-> ilAssQuestionList::load) built one large query with 7 correlated EXISTS
subqueries (feedback x4, hints, taxonomies) as SELECT fields. These were
evaluated for every candidate row BEFORE ORDER BY/LIMIT, so on large
instances (bug report: ~199k rows) the query ran 30-40 min and blocked
other queries.
Fix: when a Range is set and neither a HAVING filter nor an ORDER BY on a
computed column (feedback/hints/taxonomies) is active, load in two phases:
Phase A: SELECT question_id + required JOINs/filters + GROUP BY +
ORDER BY + LIMIT (no EXISTS subqueries) -> small paginated id set
Phase B: full SELECT incl. feedback/hints/taxonomies EXISTS, restricted
to the paginated ids via IN (...) -> flags computed only for the
~800 visible rows instead of the full candidate set.
Fallback to the original single-phase query for: no range, HAVING filter
(feedback/hints = true|false), ORDER BY feedback/hints/taxonomies.
buildOrderQueryExpression now qualifies columns for phase A (qpl_questions.*
not selected -> ambiguous title) and backticks qualified names per segment
(`qpl_questions`.`title`).
Adds composite index i6 (obj_fi, original_id, title) on qpl_questions
via ilTestQuestionPool10DBUpdateSteps::step_3() to support the phase A
filter (obj_fi IN ... AND original_id IS NULL) + default title ordering.
Adds ilAssQuestionListTwoPhaseTest (15 tests) covering the two-phase path,
all fallback conditions, column qualification and SQL shape.
Verified with EXPLAIN against the local docker DB: phase A has no
DEPENDENT SUBQUERY (3 simple JOINs only), phase B's subqueries are
evaluated over rows=2 (paginated set) instead of the full candidate set.
|
Hi @mglaubitz Thank you for the PR. First one question:
What happens if you only add the first of the two changes (the index) provided in this PR: This index seems to be added to improve the existing query (and potentially the new one) as the corresponding fields seem to not show up elsewhere in your changes. It should already improve the situation somewhat. What was the improvement if only this change was applied (see)? Questions:
Change Requests:
This is just the first round, as far as I can see. I.e. function naming seems to be somewhat lacking, but the correct names will only become clear once we really have a final structure. Once this is done the code should have become more readable and logical, so that we can figure out, what else needs to be done. Best, |
…review - Single load() path: step 1 always fetches all readily-available columns + filter/order/limit; step 2 enriches only the small paginated id set with the expensive EXISTS subqueries (feedback/ hints/taxonomies) when they are not already needed for filtering or ordering. - Remove separate canUseTwoPhaseQuery()/loadTwoPhase() execution path; range === null is no longer disallowed. - qualifyField() qualifies all fields (not only order fields); backtickField() extracted as separate helper. - No nested function calls on one line; variables in strings always wrapped in curly braces; reduced complexity. - Rewrite ilAssQuestionListTwoPhaseTest for the new structure.
…Query()
Direct string concatenation dropped the separator between
'WHERE qpl_questions.tstamp > 0' and the conditional filter
expression ('AND ...'), producing invalid SQL like '> 0AND'.
Revert to implode(PHP_EOL, array_filter([...])).
|
Thanks @kergomard for the detailed review — very helpful. We have restructured the code accordingly and pushed an updated version. Your questions:
Your change requests:
|
|
n general: I see this PR as an attempt to solve a problem / bug. The goal is not to save money! I am more than happy to fund a code review and I have already contacted a service provider that might still have some free capacity this year. In addition, I forgot something in my last post:
All the best, Marko |
kergomard
left a comment
There was a problem hiding this comment.
Hi @mglaubitz
Ok, we approaching what I would expect. First a few things:
- Thanks for answering my questions.
- Please add a "WIP" to the PR, if you feel like somebody else should go over a PR or even better make it against their repository, this way it doesn't get on our radar. I only looked at this because I had some funding and I thought the issue was interesting, but I would probably have closed it, if I wouldn't have had the resources, as this clutters our backlog. So again: for the future, make sure you do ask a programmer to provide the PR, if it changes code (clearly, documentation changes are fine).
Now, I think the general structure is starting to become more logical and understandable, I didn't go through the whole thing again, as I think there is one more improvement we should apply:
- You now added a blanket
$with_computed, I do not think this is the most efficient way: We only need to add the fields we really need, e.g. if we are neither ordering nor filtering by hints we do not need this, but we might still need the feedback sub-query as we are either ordering or filtering by that.
Some more peanuts:
- I do not like intermediate variables (I know, this is a personal preference, but it is one I hold to) and I would like you to treat variables as immutable. So, e.g. these lines should be merged.
- Please really use class-constants for the table-names. Yes, I know, we do not do this everywhere, but there is a good reason to do so: It makes it easy to find stupid typos as they can only happen in one place and table-names become very easy to change. We started using this pattern in the test in repositories. It should then clearly not be
TABLE_NAME, butQUESTION_TABLE_NAMEand so on. - Please do not check for truth, where there is none (e.g. here). Only
bools convey truth. I know, this needs more thinking about what you expect, but it makes clear what you are looking for. - I would think that the failing test is due to the new testing class you introduced.
- I assume the
PERFORMANCE_PLAN.mdwill then simply leave us in the final version.
Best,
@kergomard
…nts, review fixes - Replace blanket $with_computed boolean with per-field set: only the computed columns (feedback/hints/taxonomies) actually needed for HAVING filtering or ORDER BY are included in step 1; step 2 enriches only the remaining ones. - Add class constants for all table names (QUESTION_TABLE_NAME, etc.) and use them consistently throughout the query builders. - Fix truthiness checks: use !== null instead of implicit bool cast for nullable int properties (parentObjId, answerStatusActiveId). - Merge qualifyField()+backtickField() into qualifyAndBacktickField() to avoid intermediate variable reassignment. - Update tests for the new granular structure (22 tests).
we let one of our local AI models (GLM 5.2) analyse the problem and create a bug fix. this reduces loading time on ourt exam-test server from 3.6 minutes to 6.2 seconds (!)