{"_canonicalization":{"envelope_id":"axm_ + sha256(envelope minus {signature, axiom_id, anchors})","envelope_signature":"ed25519(envelope minus {signature, axiom_id})","json":"sort_keys=True, separators=(',',':'), ensure_ascii=False, allow_nan=False, utf-8","leaf_hash":"sha256(0x00 || canonical_json(envelope_full))","seal_signature":"ed25519(seal minus {signature, sig_algorithm})"},"axiom_id":"axm_ccce67134877427c05d72f0f8950c9aa604858eacc82331fddb3c7bc025db544","bitcoin_anchor":{"bitcoin_attestations":[],"calendar_attestations":[],"ots_url":"","stamped_at":"","status":"pending_next_stamp"},"envelope":{"anchors":[{"chain":"crovia.axiom_graph","height":0,"merkle_proof":"spider_vendor_press_v1","root_at_anchor":"spider_vendor_press_v1"}],"axiom_id":"axm_ccce67134877427c05d72f0f8950c9aa604858eacc82331fddb3c7bc025db544","axiom_type":"AX.OBS","body":{"axiom_subtype":"news.vendor_press.v1","category":"news","fingerprint":"6a37b88588fb0dfb726f7ec2d668a606b862914cd458b891ef114b095f998a09","published":"Tue, 30 Jun 2026 00:00:00 -0400","receipt_hash":"6a37b88588fb0dfb726f7ec2d668a606b862914cd458b891ef114b095f998a09","schema":"spider.news.vendor_press.v1","spider":"vendor_press","spider_record":{"axiom_subtype":"news.vendor_press.v1","category":"news","decision_hint":"POSITIVE","envelope_target":"AX.OBS","fingerprint":"6a37b88588fb0dfb726f7ec2d668a606b862914cd458b891ef114b095f998a09","observed_at":"2026-06-30T04:43:04.087680Z","parent_run_hash":"74f7ab392cc702044101fe24a76a2fdad11164cd79ce725aad6c446a477e89c5","published":"Tue, 30 Jun 2026 00:00:00 -0400","runtime_version":"0.1.0","schema":"spider.news.vendor_press.v1","source_status":200,"source_url":"https://export.arxiv.org/rss/cs.AI","spider":"vendor_press","summary_excerpt":"arXiv:2605.20256v3 Announce Type: replace-cross \nAbstract: Reinforcement learning has become a cornerstone for aligning and unlocking the reasoning capabilities of large-scale models. At its core, the training loop of GRPO and its variants alternates between rollout sampling and policy update: the policy first samples rollouts from its action space, and then updates its parameters according to the advantages computed over them. Unlike supervised learning, where each gradient step is anchored to an explicit ground-truth target, the optimal gradient direction for updating model parameters in this setting is not known a priori; the high-quality rollouts drawn during the sampling stage therefore act as the implicit \"teacher\" that guides every parameter update. However, mainstream RL algorithms such as GRPO adopt a simple sampling scheme that conditions all rollouts on the same original prompt. When a task lies beyond the policy model's current capability, this sampling scheme rarely yields","title":"FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning","url":"https://arxiv.org/abs/2605.20256","vendor":"arxiv_cs_ai"},"summary":"arXiv:2605.20256v3 Announce Type: replace-cross \nAbstract: Reinforcement learning has become a cornerstone for aligning and unlocking the reasoning capabilities of large-scale models. At its core, the training loop of GRPO and its variants alternates between rollout sampling and policy update: the policy first samples rollouts from its action space, and then updates its parameters according to the advantages computed over them. Unlike supervised learning, where each gradient step is anchored to an explicit ground-truth target, the optimal gradient direction for updating model parameters in this setting is not known a priori; the high-quality rollouts drawn during the sampling stage therefore act as the implicit \"teacher\" that guides every parameter update. However, mainstream RL algorithms such as GRPO adopt a simple sampling scheme that conditions all rollouts on the same original prompt. When a task lies beyond the policy model's current capability, this sampling scheme rarely yields","title":"FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning","vendor":"arxiv_cs_ai"},"confidence":{"method":"deterministic"},"decision":"POSITIVE","issued_at":"2026-06-30T04:43:04Z","notes":"Spider vendor_press (news) news.vendor_press.v1","object":{"captured_by":"crovia.spider.vendor_press","primary_source_url":"https://arxiv.org/abs/2605.20256"},"predecessors":[],"schema":"crovia.axiom.v1","signature":"ed25519:c87c327b7e69f2b9d19a506b7c3d272a62737c249ca18e70d9e6b9d7819dd5a7de3b525af9bfb4ef3852b46a55f2a0eefed1c08f11bff308a578732d781e0c00","signer":"crovia.substrate","subject":{"observed_at":"2026-06-30T04:43:04Z","source_collector":"spider:vendor_press","target_id":"https://arxiv.org/abs/2605.20256"},"tsa":{"authority":"crovia.substrate.bootstrap","rfc3161_token":"{\"kind\":\"crovia.bootstrap.tsa\",\"source_jsonl\":\"/opt/crovia/spider/data/news/vendor_press_v1.jsonl\",\"source_seal_merkle_root\":\"spider_vendor_press_v1\",\"upgrade_path\":\"Sessione H \\u2014 OpenTimestamps weekly anchor\"}"},"zk_mode":"clear","zk_proof":null},"ledger":{"leaf_hash":"63a1f74fe7bba40e17dfa0d12a3d1f927f4a12e767323fc69bbded7d1f9128a1","leaf_index":265205,"ledger_path":"/opt/crovia/substrate/axiom_ledger.jsonl"},"merkle_proof":{"hash_alg":"sha256","leaf_prefix":"0x00","node_prefix":"0x01","odd_leaf_rule":"duplicate_last","path":[{"sibling":"532048bf04d877346f0936e4bb3cb975553e7526dadf7cb981e33b67928f477a","side":"left"},{"sibling":"d25501d3b452b29e5ecf5c1a6b6c35df24001cdb80e28ac2c855d2f63a76e6e7","side":"right"},{"sibling":"f0f985b3e4d14f83fdd57568b05067b5e0f02b72325910bba50267dda8d6af6b","side":"left"},{"sibling":"0f687ab92c37995fd33fed61109a5c23879909ce13184b78c416a8406bdc02c6","side":"right"},{"sibling":"d7d9c084bcdb2899aa3115a2f475aa0ec5d756c37a157a8527db22fcda27d6bb","side":"left"},{"sibling":"1cc27226de3758c1c362adab110f1588ea4b17bdb628e4bbb52505c3f0e92fba","side":"left"},{"sibling":"5289e8ce54e43e43a6993c090a3cf97c23f76dc4d3199096b65fcfa71cb4cf7e","side":"left"},{"sibling":"10a48a93a0ff0e4d8f47e9a90582ecaf28e7ec0eb7e47a52bf101e3640d950f9","side":"left"},{"sibling":"3f337e26f1a0c4c64b8b7ebac04427325a7e7838c802c162576de1cd617b91ce","side":"left"},{"sibling":"1549a8883ab3267f958dc2624919e40c65c82958b8005967ed6a4da1247da0ad","side":"left"},{"sibling":"23994bf0974e5c9c7f63a61b4f0a48b0ca756a4adc34a8f85f878e774c37dfbe","side":"right"},{"sibling":"f9b4bed84fa6990c71ad2887c91bda183001f05f6b648f21d1273045a26b11fd","side":"left"},{"sibling":"173d2dc4b29ee04ea41d6d0ebc334c4bc2d46e7ee4230c94765413f24fb4bc42","side":"right"},{"sibling":"112461f7c0ec411116fb5c6c90fe95cea9d8f188b9fe08afe25a837ac02d0071","side":"right"},{"sibling":"ea9488204352c49db8f7daf05eefcd7628ecf9413830346674801a99d0654a94","side":"right"},{"sibling":"6261c13b9922cb657f10d1e5d36ec15d8771cf8766e36c61dcbffb7bed57e396","side":"right"},{"sibling":"fa19aa3faf287618b820bcfceebb366152ad521dd20ef9f51e977816663e448b","side":"right"},{"sibling":"c32f943406b62d1fc59b7f7e243492174c8e1caba8c8a2705f86c773315736e0","side":"right"},{"sibling":"1cecb7f447febd025aac272837c80de218aecc6485d2395a509b2a1f1b9c746e","side":"left"}]},"schema":"crovia.axiom_proof.v1","seal":{"first_collector_run_id":"","first_receipt_hash":"","jsonl_path":"/opt/crovia/substrate/axiom_ledger.jsonl","key_id":"430895f101d38164","last_collector_run_id":"","last_receipt_hash":"","leaf_count":265374,"merkle_root":"9636001ecab173cb6af10dc7c71eb14585daa62f9c0a6f027046f05633156891","public_key_hex":"cf742e26f75669dc673cb5c0786a1ae23ae8ca19c347317192ce40c28a7ff25c","run_id":"hourly_json_retrofit_20260630T053701Z","schema":"crovia.seal.v1","seal_family_version":"crovia-seal-family/1","seal_kind":"substrate_batch","sealed_at":"2026-06-30T05:38:03Z","sig_algorithm":"ed25519","signature":"6d4b8fd9b9da856cbb5fba7540877c6a63fa18a5ec3eaf28dc4d6d1c64921c9c0565f8c95f4b7aec0e7744fd7754844f051baf863db708cad768765b416a7b0c","signer_version":"1.1.0"},"trust_root":{"key_id":"430895f101d38164","public_key_hex":"cf742e26f75669dc673cb5c0786a1ae23ae8ca19c347317192ce40c28a7ff25c","signature_algorithm":"ed25519","url":"/registry/canon/TRUST_ROOT.md"},"verifier":{"spec":"/registry/canon/AXIOM_RECEIPT_v1.md","url":"/v/axm_ccce67134877427c05d72f0f8950c9aa604858eacc82331fddb3c7bc025db544"}}