nimbus-eth1

Commit Graph

Author	SHA1	Message	Date
Jordan Hrycaj	5b6ccddaa0	Db folder sources and related remove compiler warnings (#2673 ) * Aristo: Rename `Hash256` -> `Hash32` * CoreDb: Rename `Hash256` -> `Hash32` * Ledger: Rename `Hash256` -> `Hash32` * StorageTypes: Rename `Hash256` -> `Hash32` * Aristo: Rename `Blob` -> `seq[byte]`, `keccakHash` -> `keccak256` * Kvt: Rename `Blob` -> `seq[byte]` * CoreDb: Rename `Blob` -> `seq[byte]`, `keccakHash` -> `keccak256` * Ledger: Rename `Blob` -> `seq[byte]`, `keccakHash` -> `keccak256` * CoreDb: Rename `BlockHeader` -> `Header`, `BlockNonce` -> `Bytes8` * Misc: Rename `StorageKey` -> `Bytes32` * Tracer: `Hash256` -> `Hash32`, `BlockHeader` -> `Header`, etc. * Fix copyright header	2024-10-01 21:03:10 +00:00
Jacek Sieka	c210885b73	eth: bump to new types (#2660 ) This is a minimal set of changes to make things work with the new types in nim-eth - this is the minimal PR that merely resolves incompatibilities while the full change set would include more cleanup and migration.	2024-09-29 14:37:09 +02:00
Jacek Sieka	f3e3c6bbe0	init style for Hash256 (#2661 ) * init style for Hash256 https://github.com/status-im/nim-eth/pull/733 updates `Hash256` to become an array instead of an object - unfortunately, nim does not allow constructing arrays with `name()`, so this PR changes it to `default` which works with both. * lint	2024-09-26 13:24:36 +02:00
Jacek Sieka	513f11f911	bumps (#2652 ) eth/stew/unittest2 in preparation for eth refactoring	2024-09-24 13:19:09 +02:00
Jacek Sieka	7a15aa2a3a	clean up vertex delete (#2644 ) avoid allocating and updating the trie twice when the branch is fully removed	2024-09-20 10:31:29 +02:00
Jacek Sieka	b4b4d16729	speed up key computation (#2642 ) * batch database key writes during `computeKey` calls * log progress when there are many keys to update * avoid evicting the vertex cache when traversing the trie for key computation purposes * avoid storing trivial leaf hashes that directly can be loaded from the vertex	2024-09-20 07:43:53 +02:00
Jacek Sieka	2fe8cc4551	leaf cache fixes (#2637 ) * Add missing leaf cache update when a leaf turns to a branch with two leaves (on merge) and vice versa (on delete) - this could lead to stale leaves being returned from the cache causing validation failures - it didn't happen because the leaf caches were not being used efficiently :) * Replace `seq` with `ArrayBuf` in `Hike` allowing it to become allocation-free - this PR also works around an inefficiency in nim in returning large types via a `var` parameter * Use the leaf cache instead of `getVtxRc` to fetch recent leaves - this makes the vertex cache more efficient at caching branches because fewer leaf requests pass through it.	2024-09-19 10:39:06 +02:00
Jacek Sieka	5cd0297462	fix missed cache opportunity (#2628 ) The storage leaf cache was being circumvented when actually fetching leaves and was instead only being filled with items :/ Also avoids an expensive copy when fetching account data (broadly, variant objects are comparatively expensive to copy and fetching accounts is a hotspot)	2024-09-14 09:47:32 +02:00
Jacek Sieka	adb8d64377	simplify VertexRef (#2626 ) * move pfx out of variant which avoids pointless field type panic checks and copies on access * make `VertexRef` a non-inheritable object which reduces its memory footprint and simplifies its use - it's also unclear from a semantic point of view why inheritance makes sense for storing keys	2024-09-13 18:55:17 +02:00
Jacek Sieka	5c1e2e7d3b	Migrate `keyed_queue` to `minilru` (#2608 ) Compared to `keyed_queue`, `minilru` uses significantly less memory, in particular for the 32-byte hash keys where `kq` stores several copies of the key redundantly.	2024-09-13 15:47:50 +02:00
Jordan Hrycaj	c6674311eb	Fringe case portal proof for existing account without storage tree (#2613 ) detail: For practical reasons, ifsuch an account is asked for a slot, an empty proof list is returned. It is up to the user to provide an account proof that shows that there is no storage tree.	2024-09-11 20:27:42 +00:00
Jordan Hrycaj	75808bc03b	Add portal proof functionality for non-existing keys/paths (#2610 )	2024-09-11 09:39:45 +00:00
tersec	1b173d420d	small cleanups (#2598 ) * small cleanups * stop hiding ConvFromXtoItselfNotNeeded hints * lowmem optimization flag is no-op	2024-09-10 05:24:45 +00:00
Jacek Sieka	d39c589ec3	lru cache updates (#2590 ) * replace rocksdb row cache with larger rdb lru caches - these serve the same purpose but are more efficient because they skips serialization, locking and rocksdb layering * don't append fresh items to cache - this has the effect of evicting the existing items and replacing them with low-value entries that might never be read - during write-heavy periods of processing, the newly-added entries were evicted during the store loop * allow tuning rdb lru size at runtime * add (hidden) option to print lru stats at exit (replacing the compile-time flag) pre: ``` INF 2024-09-03 15:07:01.136+02:00 Imported blocks blockNumber=20012001 blocks=12000 importedSlot=9216851 txs=1837042 mgas=181911.265 bps=11.675 tps=1870.397 mgps=176.819 avgBps=10.288 avgTps=1574.889 avgMGps=155.952 elapsed=19m26s458ms ``` post: ``` INF 2024-09-03 13:54:26.730+02:00 Imported blocks blockNumber=20012001 blocks=12000 importedSlot=9216851 txs=1837042 mgas=181911.265 bps=11.637 tps=1864.384 mgps=176.250 avgBps=11.202 avgTps=1714.920 avgMGps=169.818 elapsed=17m51s211ms ``` 9%:ish import perf improvement on similar mem usage :)	2024-09-05 11:18:32 +02:00
Jordan Hrycaj	3c6400673d	Coredb fix kvt save only fringe condition (#2592 ) * Cosmetics, spelling, etc. * Aristo: make sure that a save cycle always commits even when empty why: If `Kvt` is tied to the `Aristo` DB save cycle, then this save cycle must also be committed if there is no data to save for `Aristo`. Otherwise this will lead to excessive core memory use with some fringe condition where Eth headers (or blocks) are downloaded while syncing and not really stored on disk. * CoreDb: Correct persistent save mode why: Saving `Kvt` first is seen as a harbinger (or canary) for `Aristo` as both run in sync. If `Kvt` succeeds saving first, so must be `Aristo` next. Other than this is a defect.	2024-09-04 13:48:38 +00:00
Jacek Sieka	35cc78c86d	add metrics for rdb lru cache (#2586 ) This is a first step towards measuring the efficiency of the LRU caches over time - metrics can be collected during import or when running regulary. Since `nim-metrics` carries some overhead for its default way of reporting metrics, this PR implements a custom collector over atomic counters, given that this is one of the hottest spots in the block processing pipeline. Using a compile-time flag, the same metrics can be printed on exit which is useful when comparing different strategies for caching - here's a recent run over blocks 16000001-1616384 - this is a good candidate to expose in a better way in the future, maybe: ``` state vtype miss hit total hitrate Account Leaf 4909417 4466215 9375632 47.64% Account Branch 20742574 72015123 92757697 77.64% World Leaf 940483 1140946 2081429 54.82% World Branch 8224151 131496580 139720731 94.11% all all 34816625 209118864 243935489 85.73% ```	2024-09-02 17:34:10 +02:00
Jacek Sieka	ef1bab0802	avoid some trivial memory allocations (#2587 ) * pre-allocate `blobify` data and remove redundant error handling (cannot fail on correct data) * use threadvar for temporary storage when decoding rdb, avoiding closure env * speed up database walkers by avoiding many temporaries ~5% perf improvement on block import, 100x on database iteration (useful for building analysis tooling)	2024-09-02 16:03:10 +02:00
Jordan Hrycaj	a25ea63dec	Revert lazy implementation (#2585 )	2024-09-02 10:34:42 +00:00
Jacek Sieka	5941fef211	Avoid unnecessary layer copy (#2567 ) When the stack has an empty layer on top, there's no need to copy the contents of `top` to it since it would be the same. ~13% processing saved (!) pre ``` INF 2024-08-17 19:11:31.748+02:00 Imported blocks blockNumber=18667648 blocks=12000 importedSlot=7860043 txs=1797812 mgas=181135.177 bps=8.763 tps=1375.062 mgps=132.125 avgBps=6.798 avgTps=1018.501 avgMGps=102.617 elapsed=29m25s154ms ``` post ``` INF 2024-08-17 18:22:52.513+02:00 Imported blocks blockNumber=18667648 blocks=12000 importedSlot=7860043 txs=1797812 mgas=181135.177 bps=9.648 tps=1513.961 mgps=145.472 avgBps=7.876 avgTps=1179.998 avgMGps=118.888 elapsed=25m23s572ms ```	2024-08-19 08:46:04 +02:00
Jordan Hrycaj	cbe5131927	Simplify aristo tree deletion functionality (#2563 ) * Cleaning up, removing cruft and debugging statements * Make `aristo_delta` fluffy compatible why: A sub-module that uses `chronicles` must import all possible modules used by a parent module that imports the sub-module. * update TODO	2024-08-14 12:09:30 +00:00
Jordan Hrycaj	ce713d95fc	Aristo lazily delete larger subtrees (#2560 ) * Extract sub-tree deletion functions into separate sub-modules * Move/rename `aristo_desc.accLruSize` => `aristo_constants.ACC_LRU_SIZE` * Lazily delete sub-trees why: This gives some control of the memory used to keep the deleted vertices in the cached layers. For larger sub-trees, keys and vertices might be on the persistent backend to a large extend. This would pull an amount of extra information from the backend into the cached layer. For lazy deleting it is enough to remember sub-trees by a small set of (at most 16) sub-roots to be processed when storing persistent data. Marking the tree root deleted immediately allows to let most of the code base work as before. * Comments and cosmetics * No need to import all for `Aristo` here * Kludge to make `chronicle` usage in sub-modules work with `fluffy` why: That `fluffy` would not run with any logging in `core_deb` is a problem I have known for a while. Up to now, logging was only used for debugging. With the current `Aristo` PR, there are cases where logging might be wanted but this works only if `chronicles` runs without the `json[dynamic]` sinks. So this should be re-visited. * More of a kludge	2024-08-14 08:54:44 +00:00
Jordan Hrycaj	7becf4e389	Remove vertex ID recycle function (#2558 ) why: It is not safe in general to recycle vertex IDs while the `RocksDb` cache has `VertexID` rather than `RootedVertexID` where the former type seems preferable. In some fringe cases one might remove a vertex with key `(root1,vid)` and insert another vertex with key `(root2,vid)` while re-using the vertex ID `vid`. Without knowledge of `root1` and `root2`, the LRU cache will return the same vertex for `(root2,vid)` also for `(root1,vid)`.	2024-08-12 20:56:15 +00:00
Jacek Sieka	094486d0ce	Hash bump	2024-08-08 07:46:35 +02:00
Jordan Hrycaj	38572bd8ea	Cache a storage root ID forever in the leaf payload of an account (#2551 ) details: Stale root IDs are marked disabled while the ID is kept in the leaf payload. why: This might lead to further caching advantages.	2024-08-07 13:28:01 +00:00
Jordan Hrycaj	488bdbc267	Provide portal proof functionality with coredb (#2550 ) * Provide portal proof functions in `aristo_api` why: So it can be fully supported by `CoreDb` * Fix prototype in `kvt_api` * Fix node constructor for account leafs with storage trees * Provide simple path check based on portal proof functionality * Provide portal proof functionality in `CoreDb` * Update TODO list	2024-08-07 11:30:55 +00:00
Jordan Hrycaj	6bae929439	Added comments (#2546 )	2024-08-06 12:43:39 +00:00
Jordan Hrycaj	5b502a06c4	Added portal proof nodes generation functionality (#2539 ) * Extracted `test_tx.testTxMergeProofAndKvpList()` => separate file * Fix serialiser why: Typo lead to duplicate rlp-encoded nodes in chain * Remove cruft * Implemnt portal proof nodes generators `partXxxTwig()` * Add unit test for portal proof nodes generator `partAccountTwig()` * Cosmetics * Simplify serialiser return code format * Fix proof generator for extension nodes why: Code was simply bonkers, not detected before the unit tests were adapted to check for just this. * Implemented portal proof nodes verifier `partUntwig()` * Cosmetics * Fix `testutp` cli poblem	2024-08-06 11:29:26 +00:00
Jordan Hrycaj	01b5c08763	Revive json tracer unit tests (#2538 ) * Some `Aristo` clean-ups/updates * Re-implemented core-db tracer functionality * Rename nimbus tracer `no-tracer.nim` => `tracer.nim` why: Restore original name for easy diff tracking with upcoming update * Update nimbus tracer using new core-db tracer functionality * Updating json tracer unit tests * Enable json tracer unit tests	2024-08-01 10:41:20 +00:00
Jordan Hrycaj	72c3ab8ced	Provide partial tree support for preloading tests (#2536 ) * Implement partial trees why: This is currently needed for unit tests to pre-load the database with test data similar to `proof` node pre-load. The basic features for `snap-sync` boundary proofs are available as well for future use. What is missing is the final proof verification and a complete storage data load/merge function (stub is available.) * Cosmetics, clean up	2024-07-29 20:15:17 +00:00
andri lim	254bda365f	Remove txpool sender locality (#2525 ) * Remove txpool sender locality We no longer distinct local or remote sender * Fix copyright year	2024-07-25 22:36:08 +07:00
Jordan Hrycaj	1452e7b1c0	Misc updates (#2513 ) * Update config for Ledger and CoreDb why: Prepare for tracer which depends on the API jump table (as well as the profiler.) The API jump table is now enabled in unit/integration test mode piggybacking on the `unittest2DisableParamFiltering` compiler flag or on an extra compiler flag `dbjapi_enabled`. * No deed for error field in `NodeRef` why: Was opnly needed by proof nodes pre-loader which will be re-implemented * Cosmetics	2024-07-22 18:10:04 +00:00
Jordan Hrycaj	5ac362fe6f	Aristo and kvt balancer management update (#2504 ) * Aristo: Merge `delta_siblings` module into `deltaPersistent()` * Aristo: Add `isEmpty()` for canonical checking whether a layer is empty * Aristo: Merge `LayerDeltaRef` into `LayerObj` why: No need to maintain nested object refs anymore. Previously the `LayerDeltaRef` object had a companion `LayerFinalRef` which held non-delta layer information. * Kvt: Merge `LayerDeltaRef` into `LayerRef` why: No need to maintain nested object refs (as with `Aristo`) * Kvt: Re-write balancer logic similar to `Aristo` why: Although `Kvt` was a cheap copy of `Aristo` it sort of got out of sync and the balancer code was wrong. * Update iterator over forked peers why: Yield additional field `isLast` indicating that the last iteration cycle was approached. * Optimise balancer calculation. why: One can often avoid providing a new object containing the merge of two layers for the balancer. This avoids copying tables. In some cases this is replaced by `hasKey()` look ups though. One uses one of the two to combine and merges the other into the first. Of course, this needs some checks for making sure that none of the components to merge is eventually shared with something else. * Fix copyright year	2024-07-18 21:32:32 +00:00
Jacek Sieka	df4a21c910	Store cached hash at the layer corresponding to the source data (#2492 ) When lazily verifying state roots, we may end up with an entire state without roots that gets computed for the whole database - in the current design, that would result in hashes for the entire trie being held in memory. Since the hash depends only on the data in the vertex, we can store it directly at the top-most level derived from the verticies it depends on - be that memory or database - this makes the memory usage broadly linear with respect to the already-existing in-memory change set stored in the layers. It also ensures that if we have multiple forks in memory, hashes get cached in the correct layer maximising reuse between forks. The same layer numbering scheme as elsewhere is reused, where -2 is the backend, -1 is the balancer, then 0+ is the top of the stack and stack. A downside of this approach is that we create many small batches - a future improvement could be to collect all such writes in a single batch, though the memory profile of this approach should be examined first (where is the batch kept, exactly?).	2024-07-18 09:13:56 +02:00
Jordan Hrycaj	6677f57ea9	Aristo balancer clean up (#2501 ) * Remove `chunkedMpt` from `persistent()`/`stow()` function why: Proof-mode code was removed with PR #2445 and needs to be re-designed. * Remove unused `beStateRoot` argument from `deltaMerge()` * Update/drastically simplify `txStow()` why: Got rid of many boundary conditions details: Many pre-conditions have changed. In particular, previous versions used the account state (hash) which was conveniently available and checked it against the backend in order to find out whether there was something to do, at all. Currently, only an empty set of all tables in the delta layer has the balancer update ignored. Notable changes are: * no check against account state (see above) * balancer filters have no hash signature (some legacy stuff left over from journals) * no (shap sync) proof data which made the generation of the a top layer more complex * Cosmetics, cruft removal * Update unit test file & function name why: Was legacy module	2024-07-17 19:27:33 +00:00
Jordan Hrycaj	17391b58d0	Hash keys and hash256 revisited (#2497 ) * Remove cruft left-over from PR #2494 * TODO * Update comments on `HashKey` type values * Remove obsolete hash key conversion flag `forceRoot` why: Is treated implicitly by having vertex keys as `HashKey` type and root vertex states converted to `Hash256`	2024-07-17 20:48:21 +07:00
andri lim	a59cc84fca	Not using deprecated functions in config anymore (#2495 )	2024-07-17 02:57:19 +00:00
Jordan Hrycaj	a84a2131cd	No ext update (#2494 ) * Imported/rebase from `no-ext`, PR #2485 Store extension nodes together with the branch Extension nodes must be followed by a branch - as such, it makes sense to store the two together both in the database and in memory: * fewer reads, writes and updates to traverse the tree * simpler logic for maintaining the node structure * less space used, both memory and storage, because there are fewer nodes overall There is also a downside: hashes can no longer be cached for an extension - instead, only the extension+branch hash can be cached - this seems like a fine tradeoff since computing it should be fast. TODO: fix commented code * Fix merge functions and `toNode()` * Update `merkleSignCommit()` prototype why: Result is always a 32bit hash * Update short Merkle hash key generation details: Ethereum reference MPTs use Keccak hashes as node links if the size of an RLP encoded node is at least 32 bytes. Otherwise, the RLP encoded node value is used as a pseudo node link (rather than a hash.) This is specified in the yellow paper, appendix D. Different to the `Aristo` implementation, the reference MPT would not store such a node on the key-value database. Rather the RLP encoded node value is stored instead of a node link in a parent node is stored as a node link on the parent database. Only for the root hash, the top level node is always referred to by the hash. * Fix/update `Extension` sections why: Were commented out after removal of a dedicated `Extension` type which left the system disfunctional. * Clean up unused error codes * Update unit tests * Update docu --------- Co-authored-by: Jacek Sieka <jacek@status.im>	2024-07-16 19:47:59 +00:00
Jacek Sieka	9d91191154	storage hike cache (#2484 ) This PR adds a storage hike cache similar to the account hike cache already present - this cache is less efficient because account storage is already partically cached in the account ledger but nonetheless helps keep hiking down. Notably, there's an opportunity to optimise this cache and the others so that they cooperate better insteado of overlapping, which is left for a future PR. This PR also fixes an O(N) memory usage for storage slots where the delete would keep the full storage in a work list which on mainnet can grow very large - the work list is replaced with a more conventional recursive `O(log N)` approach.	2024-07-14 19:12:10 +02:00
Jacek Sieka	f3a56002ca	Turn payload into value type (#2483 ) The Vertex type unifies branches, extensions and leaves into a single memory area where the larges member is the branch (128 bytes + overhead) - the payloads we have are all smaller than 128 thus wrapping them in an extra layer of `ref` is wasteful from a memory usage perspective. Further, the ref:s must be visited during the M&S phase of garbage collection - since we keep millions of these, many of them short-lived, this takes up significant CPU time. ``` Function CPU Time: Total CPU Time: Self Module Function (Full) Source File Start Address system::markStackAndRegisters 10.0% 4.922s nimbus system::markStackAndRegisters(var<system::GcHeap>).constprop.0 gc.nim 0x701230` ```	2024-07-14 12:02:05 +02:00
Jacek Sieka	72947b3647	odds and ends (#2481 ) small cleanups to reduce memory allocations	2024-07-13 20:42:49 +02:00
Jordan Hrycaj	b924fdcaa7	Separate config for core db and ledger (#2479 ) * Updates and corrections * Extract `CoreDb` configuration from `base.nim` into separate module why: This makes it easier to avoid circular imports, in particular when the capture journal (aka tracer) is revived. * Extract `Ledger` configuration from `base.nim` into separate module why: This makes it easier to avoid circular imports (if any.) also: Move `accounts_ledger.nim` file to sub-folder `backend`. That way the layout resembles that of the `core_db`.	2024-07-12 13:12:25 +00:00
Jacek Sieka	01ab209497	cache account payload (#2478 ) Instead of caching just the storage id, we can cache the full payload which further reduces expensive hikes	2024-07-12 15:08:26 +02:00
Jacek Sieka	a6764670f0	merge: avoid hike allocations (#2472 ) hike allocations (and the garbage collection maintenance that follows) are responsible for some 10% of cpu time (not wall time!) at this point - this PR avoids them by stepping through the layers one step at a time, simplifying the code at the same time.	2024-07-11 13:26:46 +02:00
Jordan Hrycaj	800fd77333	Core db remove legacy phrases (#2468 ) * Rename `newKvt()` -> `ctx.getKvt()` why: Clean up legacy shortcut. Also, the `KVT` returned is not instantiated but refers to the shared `KVT` that resides in a context which is a generalisation of an in-memory database fork. The function `ctx` retrieves the default context. * Rename `newTransaction()` -> `ctx.newTransaction()` why: Clean up legacy shortcut. The transaction is applied to a context as a generalisation of an in-memory database fork. The function `ctx` retrieves the default context. * Rename `getColumn(CtGeneric)` -> `getGeneric()` why: No more a list of well known sub-tries needed, a single one is enough. In fact, `getColumn()` did only support a single sub-tree by now. * Reduce TODO list	2024-07-10 12:19:35 +00:00
Jacek Sieka	3382c2427b	increase rdb cache sizes (#2466 ) This trivial bump should improve performance a bit without costing too much memory - as the trie grows, so does the number of levels in it and creating hikes becomes ever more expensive - hopefully this cache increase should give a nice little boost even if it's not a lot.	2024-07-09 17:35:27 +02:00
andri lim	c775c906a2	Fix LedgerRef storage iterator and add test (#2458 )	2024-07-05 10:15:48 +00:00
Jacek Sieka	7d78fd97d5	avoid allocations for slot storage (#2455 ) Introduce a new `StoData` payload type similar to `AccountData` * slightly more efficient storage format * typed api * fewer seqs * fix encoding docs - it wasn't rlp after all :)	2024-07-04 23:48:45 +00:00
Jacek Sieka	81e75622cf	storage: store root id together with vid, for better locality of refe… (#2449 ) The state and account MPT:s currenty share key space in the database based on that vertex id:s are assigned essentially randomly, which means that when two adjacent slot values from the same contract are accessed, they might reside at large distance from each other. Here, we prefix each vertex id by its root causing them to be sorted together thus bringing all data belonging to a particular contract closer together - the same effect also happens for the main state MPT whose nodes now end up clustered together more tightly. In the future, the prefix given to the storage keys can also be used to perform range operations such as reading all the storage at once and/or deleting an account with a batch operation. Notably, parts of the API already supported this rooting concept while parts didn't - this PR makes the API consistent by always working with a root+vid.	2024-07-04 15:46:52 +02:00
Jacek Sieka	b23795ab39	remove pPrf, fRpp (#2445 ) No longer used now that hashify is gone	2024-07-03 22:21:57 +02:00
Jacek Sieka	443c6d1f8e	Cache account path storage id (#2443 ) The storage id is frequently accessed when executing contract code and finding the path via the database requires several hops making the process slow - here, we add a cache to keep the most recently used account storage id:s in memory. A possible future improvement would be to cache all account accesses so that for example updating the balance doesn't cause several hikes.	2024-07-03 17:58:25 +02:00

1 2 3 4

179 Commits