summaryrefslogtreecommitdiff
path: root/Documentation
diff options
context:
space:
mode:
authorPaulo Alcantara <pc@manguebit.org>2026-09-21 11:17:26 -0300
committerPaulo Alcantara <pc@manguebit.org>2026-09-30 14:22:44 -0300
commit69bf14300b50631e7c2271289d347bbd1b19a06c (patch)
treec3b3dbc1390ef02f0521236d243476aeb55e28e7 /Documentation
parentf14572c203d57492e1d4e5d7851a3b143e083b82 (diff)
netfs: clear post-EOF pagecache when extending a file via write
Fix netfs to erase the contents of a hole created after the EOF by an ordinary write if dirty data has been previously left there by writes through an mmapped region. Neither the buffered nor the unbuffered/DIO write path clears that stale pagecache. Zero the tail of the folio straddling the EOF before an extending write. That is the only folio that can hold data written past the EOF through an mmap, as pages wholly beyond the EOF can't be faulted in. The folio is zeroed rather than dropped so a concurrent extending write can't lose data. Both write paths downgrade the i_rwsem to shared, so extending writes can run concurrently and the i_size read by the caller may be stale by the time the folio is locked. Re-read i_size under the folio lock and clamp the zeroed range up to it, so a racing write that already put data into the folio isn't clobbered. Wait for any writeback on the folio to finish before zeroing it so that the pagecache isn't modified while it may still be read by the transport during transmission. Honour IOCB_NOWAIT by returning -EAGAIN rather than blocking on the folio lock, on writeback, or in folio_mkclean()'s rmap walk when the folio is mapped. truncate_pagecache() can't be used here: it must be called with the i_rwsem held exclusively, but these write paths only hold it shared, and it would block unconditionally, breaking IOCB_NOWAIT. Callers that hold i_rwsem exclusively for the whole resize (truncate, setattr, fallocate, clone) exclude any genuine concurrent buffered writer, so staleness can instead be decided from the folio's dirty state, as pagecache_isize_extended() already does for filesystems that serialise writes against truncate/setattr via a single i_rwsem. Export netfs_clear_stale_post_isize() helper to handle such case. The helper is required by the CIFS client to fix generic/363. Closes: https://sashiko.dev/#/patchset/20260921230755.1133425-1-pc%40manguebit.org Fixes: 938e13a73b24 ("netfs: Implement buffered write API") Fixes: 153a9961b551 ("netfs: Implement unbuffered/DIO write support") Reviewed-by: David Howells <dhowells@redhat.com> Reviewed-by: Namjae Jeon <linkinjeon@kernel.org> Signed-off-by: Paulo Alcantara <pc@manguebit.org> Cc: Christian Brauner <brauner@kernel.org> Cc: Matthew Wilcox <willy@infradead.org> Cc: Ronnie Sahlberg <ronniesahlberg@gmail.com> Cc: Shyam Prasad N <sprasad@microsoft.com> Cc: Tom Talpey <tom@talpey.com> Cc: Bharath SM <bharathsm@microsoft.com> Cc: stable@vger.kernel.org
Diffstat (limited to 'Documentation')
-rw-r--r--Documentation/filesystems/netfs_library.rst26
1 files changed, 26 insertions, 0 deletions
diff --git a/Documentation/filesystems/netfs_library.rst b/Documentation/filesystems/netfs_library.rst
index ddd799df6ce3..a9de281db8ee 100644
--- a/Documentation/filesystems/netfs_library.rst
+++ b/Documentation/filesystems/netfs_library.rst
@@ -451,6 +451,32 @@ one.
The inode should be marked ``NETFS_ICTX_SINGLE_NO_UPLOAD`` if this API is to be
used. The writeback function requires the buffer to be of ITER_FOLIOQ type.
+Clearing Stale Post-EOF Pagecache
+---------------------------------
+
+When a file is extended, data left in the pagecache past the old EOF by a write
+through an mmap must not be exposed as file content. Netfslib clears this on
+its own write paths, and exports a helper so a filesystem can do the same from
+a resize path (truncate, setattr, fallocate and the like) that holds the
+inode's ``i_rwsem`` exclusively across the whole resize::
+
+ void netfs_clear_stale_post_isize(struct inode *inode, uoff_t from,
+ uoff_t to);
+
+This zeroes any such data within the ``[from, to)`` hole to be made, where
+@from is the old EOF and @to is the new one, and the caller must have already
+updated ``i_size`` to @to before calling it. Only the folio straddling @from
+can hold data written past the EOF through an mmap, as pages wholly beyond the
+EOF can't be faulted in, so the zeroing is limited to that folio. The folio is
+zeroed rather than dropped so that a concurrent extending write can't lose
+data.
+
+This overlaps with ``pagecache_isize_extended()`` but can't reuse it: that
+helper is keyed on a sub-page block size and is a no-op when the block size is
+``>= PAGE_SIZE`` (as on network filesystems), and it doesn't wait for
+writeback. As both address the same problem, a change to one should probably
+be reflected in the other to keep them in sync.
+
High-Level VM API
==================