Repository navigation
Add a dominator-tree WTO utility for reducible CFGs #9215
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
base: main
Are you sure you want to change the base?
Changes from all commits
dad6624
50f7064
08750e7
4aea7cd
a469c96
f4d0e76
f1a4061
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,321 @@ | ||
| /* | ||
| * Copyright 2026 WebAssembly Community Group participants | ||
| * | ||
| * Licensed under the Apache License, Version 2.0 (the "License"); | ||
| * you may not use this file except in compliance with the License. | ||
| * You may obtain a copy of the License at | ||
| * | ||
| * http://www.apache.org/licenses/LICENSE-2.0 | ||
| * | ||
| * Unless required by applicable law or agreed to in writing, software | ||
| * distributed under the License is distributed on an "AS IS" BASIS, | ||
| * WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. | ||
| * See the License for the specific language governing permissions and | ||
| * limitations under the License. | ||
| */ | ||
|
|
||
| // | ||
| // Weak Topological Ordering (WTO) and worklist runner for forward data flow | ||
| // analysis over reducible CFGs. | ||
| // | ||
| // A Weak Topological Ordering (Bourdoncle, "Efficient chaotic iteration | ||
| // strategies with widenings", 1993) is a hierarchical ordering of the reachable | ||
| // blocks of a directed graph in which strongly connected components (loops) are | ||
| // parenthesized into nested cycles. The first element of each cycle is its | ||
| // "head" (loop header). Formally, the WTO of a directed graph is a hierarchical | ||
| // ordering of its vertices such that for every edge u -> v, either: | ||
| // | ||
| // 1. u < v (i.e. this is a forward edge) and v is not the head of a cycle | ||
| // containing u. | ||
| // 2. u >= v (i.e. this is a backedge) and v is the head of a cycle containing | ||
| // u. | ||
| // | ||
| // Examples (writing `(h ...)` for a cycle with head `h`): | ||
| // | ||
| // - Diamond (0 -> 1, 0 -> 2, 1 -> 3, 2 -> 3): | ||
| // 0 1 2 3 | ||
| // | ||
| // - Simple loop (0 -> 1 -> 2 -> 1, 2 -> 3): | ||
| // 0 (1 2) 3 | ||
| // Here 1 is the cycle head, 1 and 2 form the cycle body, and the exit | ||
| // block 3 is outside the cycle. | ||
| // | ||
| // - Nested loops (0 -> 1 -> 2 -> 3 -> 2, 3 -> 4 -> 1, 4 -> 5): | ||
| // 0 (1 (2 3) 4) 5 | ||
| // Here outer cycle (1 (2 3) 4) with head 1 encloses inner cycle (2 3) with | ||
| // head 2. | ||
| // | ||
| // During forward dataflow analysis, elements of the WTO are evaluated | ||
| // left-to-right. When a cycle `(h ...)` is reached, its elements are evaluated | ||
| // repeatedly in order until the head `h` is no longer re-queued by a backedge. | ||
| // Inner cycles therefore stabilize completely on each iteration of an enclosing | ||
| // outer cycle before flow values propagate past the cycle. | ||
| // | ||
| // Algorithm sketch: | ||
| // | ||
| // In a reducible CFG whose blocks are ordered in reverse postorder (RPO, as | ||
| // produced by cfg-traversal.h), every cycle is a natural loop headed by a | ||
| // single entry block that dominates all blocks in the cycle, and every | ||
| // backedge `p -> h` satisfies `h` dominates `p` (with `h <= p` in RPO). We | ||
| // construct the WTO directly from the dominator tree in three steps: | ||
| // | ||
| // 1. Compute the dominator tree (`DomTree`) over the RPO-indexed blocks. | ||
| // 2. Discover natural loops from innermost to outermost by scanning candidate | ||
| // headers `h` in reverse RPO order (N - 1 down to 0). For each `h` that | ||
| // has at least one backedge `p -> h` (where `h` dominates `p`), run a | ||
| // backward DFS over predecessors starting from `p` and stopping at `h` to | ||
| // visit every block in `h`'s natural loop. Because inner loop headers have | ||
| // larger RPO indices than outer loop headers and are processed first, the | ||
| // first loop that visits a block `b != h` is its immediately enclosing | ||
| // loop (`loopParent[b] = h`). | ||
| // 3. Link each reachable block into the child list of its `loopParent` in | ||
| // increasing RPO order, then walk the resulting loop nesting forest to | ||
| // emit each loop header `h` and its children as a nested `Cycle`. | ||
| // | ||
|
|
||
| #ifndef cfg_wto_h | ||
| #define cfg_wto_h | ||
|
|
||
| #include <cassert> | ||
| #include <memory> | ||
| #include <variant> | ||
| #include <vector> | ||
|
|
||
| #include "cfg/domtree.h" | ||
| #include "wasm.h" | ||
|
|
||
| namespace wasm { | ||
|
|
||
| // The BasicBlock type is assumed to have an `in` vector of predecessor block | ||
| // pointers and a `contents.index` field of type `Index`. | ||
| template<typename BasicBlock> struct WeakTopologicalOrdering { | ||
| static constexpr Index NoIndex = Index(-1); | ||
|
|
||
| struct Cycle; | ||
| using Element = std::variant<BasicBlock*, Cycle>; | ||
| using List = std::vector<Element>; | ||
|
|
||
| struct Cycle { | ||
| List elems; | ||
|
|
||
| BasicBlock* head() const { return std::get<BasicBlock*>(elems.front()); } | ||
| bool operator==(const Cycle& other) const { return elems == other.elems; } | ||
| }; | ||
|
|
||
| List elems; | ||
|
|
||
| WeakTopologicalOrdering(std::vector<std::unique_ptr<BasicBlock>>& blocks); | ||
| }; | ||
|
|
||
| template<typename BasicBlock> | ||
| WeakTopologicalOrdering<BasicBlock>::WeakTopologicalOrdering( | ||
| std::vector<std::unique_ptr<BasicBlock>>& blocks) { | ||
| Index numBlocks = blocks.size(); | ||
| if (numBlocks == 0) { | ||
| return; | ||
| } | ||
|
|
||
| for (Index i = 0; i < numBlocks; ++i) { | ||
| blocks[i]->contents.index = i; | ||
| } | ||
|
|
||
| // TODO: Avoid building an unordered_map of block indices in DomTree when | ||
| // BasicBlock already stores its RPO index on `contents`. | ||
| DomTree<BasicBlock> domTree(blocks); | ||
|
|
||
| auto isReachable = [&](Index i) { | ||
| return i == 0 || domTree.iDoms[i] != domTree.nonsense; | ||
| }; | ||
|
|
||
| auto dominates = [&](Index dom, Index node) { | ||
| assert(isReachable(dom)); | ||
| if (!isReachable(node)) { | ||
| return false; | ||
| } | ||
| Index curr = node; | ||
| while (curr > dom) { | ||
| curr = domTree.iDoms[curr]; | ||
| } | ||
| return curr == dom; | ||
| }; | ||
|
|
||
| struct Node { | ||
| // The innermost loop header for the cycle containing this block. | ||
| Index loopParent = NoIndex; | ||
| // For loop headers, the index of their first child (i.e. the head of a | ||
| // linked list of children). | ||
| Index firstChild = NoIndex; | ||
| // A linked list edge to the next child with the same loop header. | ||
| Index nextSibling = NoIndex; | ||
| // The index of the loop header we last traversed this node for, used | ||
| // instead of a `visited` set during the DFS. | ||
| Index lastVisitedBy = NoIndex; | ||
| bool isLoopHeader = false; | ||
| }; | ||
| std::vector<Node> nodes(numBlocks); | ||
|
|
||
| // Discover natural loops from innermost to outermost (reverse RPO order). | ||
| // Because inner loops are processed before outer loops, the first loop whose | ||
| // natural loop body contains a block is its immediately enclosing loop. | ||
| // | ||
| // TODO: Collapse inner loops with union-find during natural loop discovery so | ||
| // outer loops do not re-traverse inner loop bodies. | ||
| std::vector<Index> worklist; | ||
| for (Index i = numBlocks; i > 0; --i) { | ||
| Index h = i - 1; | ||
| if (!isReachable(h)) { | ||
| continue; | ||
| } | ||
| // Check if h is the head of a loop. It is a loop header if and only if it | ||
| // dominates one of its predecessors. (We assume the CFG is reducible, so | ||
| // loop headers dominate all blocks in the loop bodies, including those that | ||
| // branch back to the header.) | ||
| nodes[h].lastVisitedBy = h; | ||
| for (auto* pred : blocks[h]->in) { | ||
| Index p = pred->contents.index; | ||
| if (dominates(h, p)) { | ||
| nodes[h].isLoopHeader = true; | ||
| // Avoid repeat traversals by setting lastVisitedBy = h on visited | ||
| // blocks. | ||
| if (nodes[p].lastVisitedBy != h) { | ||
| nodes[p].lastVisitedBy = h; | ||
| worklist.push_back(p); | ||
| } | ||
| } | ||
| } | ||
| // We've initialized the worklist with all the loop tails that branch | ||
| // directly back to the loop header. DFS from those loop tails back to the | ||
| // loop header (but no further). All the blocks we find during the DFS are | ||
| // part of the loop body. | ||
| while (!worklist.empty()) { | ||
| Index curr = worklist.back(); | ||
| worklist.pop_back(); | ||
| if (nodes[curr].loopParent == NoIndex) { | ||
| nodes[curr].loopParent = h; | ||
| } | ||
| for (auto* pred : blocks[curr]->in) { | ||
| Index p = pred->contents.index; | ||
| // The loop header has lastVisitedBy == h, so the search will stop | ||
| // there. | ||
| if (isReachable(p) && nodes[p].lastVisitedBy != h) { | ||
| assert(dominates(h, p) && "Expected reducible CFG"); | ||
| nodes[p].lastVisitedBy = h; | ||
| worklist.push_back(p); | ||
| } | ||
| } | ||
| } | ||
| } | ||
|
|
||
| // Link each reachable block into its parent loop's intrusive child list. | ||
| // Prepending in reverse RPO order yields increasing RPO order. | ||
| Index topFirstChild = NoIndex; | ||
| for (Index i = numBlocks; i > 0; --i) { | ||
| Index idx = i - 1; | ||
| if (!isReachable(idx)) { | ||
| continue; | ||
| } | ||
| Index parent = nodes[idx].loopParent; | ||
| if (parent == NoIndex) { | ||
| // Prepend to top-level list. | ||
| nodes[idx].nextSibling = topFirstChild; | ||
| topFirstChild = idx; | ||
| } else { | ||
| // Prepend to loop header's list. | ||
| nodes[idx].nextSibling = nodes[parent].firstChild; | ||
| nodes[parent].firstChild = idx; | ||
| } | ||
| } | ||
|
|
||
| // Traverse the linked lists of children, materializing them as WTO elements. | ||
| // Loop depth should be limited, so doing this recursively should be fine. If | ||
| // it ever causes an issue, we can un-recurse this. | ||
| // TODO: Flatten the WTO into a single contiguous vector of entries with cycle | ||
| // jump targets to avoid per-cycle vector allocations and recursion. | ||
| auto buildList = [&](auto& self, Index firstChild, List& out) -> void { | ||
| for (Index curr = firstChild; curr != NoIndex; | ||
| curr = nodes[curr].nextSibling) { | ||
| auto* block = blocks[curr].get(); | ||
| if (nodes[curr].isLoopHeader) { | ||
| Cycle cycle; | ||
| cycle.elems.emplace_back(block); | ||
| self(self, nodes[curr].firstChild, cycle.elems); | ||
| out.emplace_back(std::move(cycle)); | ||
| } else { | ||
| out.emplace_back(block); | ||
| } | ||
| } | ||
| }; | ||
|
|
||
| buildList(buildList, topFirstChild, elems); | ||
| } | ||
|
|
||
| // Given a CFG in reverse postorder (e.g. from cfg-traversal), run a forward | ||
| // fixed-point analysis over its basic blocks using a Weak Topological Ordering. | ||
| // | ||
| // Usage: | ||
| // 1. Construct `WTOWorklist work(cfg);` (which initializes `inQueue` and | ||
| // `index` on each block's `contents`). | ||
| // 2. Seed the initial block(s) to evaluate via `work.push(cfg.entry);`. | ||
| // 3. Call `work.run([&](BasicBlock* block) { ... });`. Inside the visitor | ||
| // callback, evaluate the transfer function for `block` and call | ||
| // `work.push(next)` for any successor whose input state changed and needs | ||
| // to be (re-)evaluated. | ||
| // | ||
| // The BasicBlock `contents` of the CFG must contain two fields: | ||
| // | ||
| // bool inQueue; // whether scheduled for visitation | ||
| // Index index; // basic block index in RPO | ||
| // | ||
| template<typename CFG> struct WTOWorklist { | ||
| using BasicBlock = typename CFG::BasicBlock; | ||
|
|
||
| CFG& cfg; | ||
|
|
||
| WTOWorklist(CFG& cfg) : cfg(cfg) { | ||
| auto& basicBlocks = cfg.basicBlocks; | ||
| for (Index i = 0; i < basicBlocks.size(); ++i) { | ||
| auto& contents = basicBlocks[i]->contents; | ||
| contents.inQueue = false; | ||
| contents.index = i; | ||
| } | ||
| } | ||
|
|
||
| void push(BasicBlock* block) { block->contents.inQueue = true; } | ||
|
|
||
| template<typename VisitFn> void run(VisitFn&& visit) { | ||
| // Iterate through each element in the current cycle's list (or the | ||
| // top-level list), which will be in reverse postorder. Visit those that are | ||
| // in the queue, which may push later elements to the queue. When there is a | ||
| // nested cycle, repeatedly visit it recursively until it stabilizes before | ||
| // continuing on. We could un-recurse this, but the loop depth is expected | ||
| // to be acceptably small. | ||
| // TODO: Fast-path initial entry singletons and CFGs without backedges | ||
| // without building DomTree or WTO, using CFGWalker::loopTops. | ||
| WeakTopologicalOrdering<BasicBlock> wto(cfg.basicBlocks); | ||
| auto evalList = | ||
| [&](auto& self, | ||
| const typename WeakTopologicalOrdering<BasicBlock>::List& list) | ||
| -> void { | ||
| for (const auto& elem : list) { | ||
| if (auto* block = std::get_if<BasicBlock*>(&elem)) { | ||
| if ((*block)->contents.inQueue) { | ||
| (*block)->contents.inQueue = false; | ||
| visit(*block); | ||
| } | ||
| } else { | ||
| const auto& cycle = | ||
| std::get<typename WeakTopologicalOrdering<BasicBlock>::Cycle>(elem); | ||
| BasicBlock* head = cycle.head(); | ||
| do { | ||
| self(self, cycle.elems); | ||
|
Member
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Can we avoid this recursion? Or is the idea that loop recursion is limited?
Member
Author
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. We probably could avoid this recursion, but loop depth should be relatively limited. I suggest we leave it unless it causes a problem for someone.
Member
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. sgtm, then perhaps a comment to mention that? |
||
| } while (head->contents.inQueue); | ||
| } | ||
| } | ||
| }; | ||
| evalList(evalList, wto.elems); | ||
| } | ||
| }; | ||
|
|
||
| } // namespace wasm | ||
|
|
||
| #endif // cfg_wto_h | ||
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
I follow this up to here. The worklist processed on line 173, however, is unclear to me. Maybe add some comments on what is happening here?
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
Added more comments!