KonstantinSeurer/mesa

Commit Graph

Author	SHA1	Message	Date
Jason Ekstrand	53bfcdeecf	intel/fs: Implement the new load/store_scratch intrinsics This commit fills in a number of different pieces: 1. We add support to brw_nir_lower_mem_access_bit_sizes to handle the new intrinsics. This involves simple plumbing work as well as a tiny bit of extra logic to always scalarize scratch intrinsics 2. Add code to brw_fs_nir.cpp to turn nir_load/store_scratch intrinsics into byte/dword scattered read/write messages which use the A32 stateless model. 3. Add code to lower_surface_logical_send to handle dword scattered messages and the A32 stateless model. Reviewed-by: Caio Marcelo de Oliveira Filho <caio.oliveira@intel.com>	2019-11-11 17:17:02 +00:00
Jason Ekstrand	e2297699de	intel/nir: Plumb devinfo through lower_mem_access_bit_sizes Reviewed-by: Caio Marcelo de Oliveira Filho <caio.oliveira@intel.com>	2019-11-11 17:17:02 +00:00
Jason Ekstrand	1dff48af05	intel/fs: refactor surface header setup Reviewed-by: Caio Marcelo de Oliveira Filho <caio.oliveira@intel.com>	2019-11-11 17:17:02 +00:00
Jason Ekstrand	a0999bc049	intel/fs: Add DWord scattered read/write opcodes Reviewed-by: Caio Marcelo de Oliveira Filho <caio.oliveira@intel.com>	2019-11-11 17:17:02 +00:00
Jason Ekstrand	83f04d80b0	intel/nir: Use nir_extract_bits in lower_mem_access_bit_sizes The new helper solves most of the annoying problems with data wrangling in brw_nir_lower_mem_access_bit_sizes. Reviewed-by: Caio Marcelo de Oliveira Filho <caio.oliveira@intel.com>	2019-11-11 17:17:02 +00:00
Jason Ekstrand	b8d45d9307	nir: Add tests for nir_extract_bits	2019-11-11 17:17:02 +00:00
Jason Ekstrand	d0bbf98c96	nir/builder: Add a nir_extract_bits helper This new helper is better than nir_bitcast_vector because it's able to take a (mostly) arbitrary range from the source vector. The only requirement is that first_bit has to be aligned to the smaller of the two bit sizes. It wouldn't be hard to lift that requirement but it's reasonable for now. Reviewed-by: Caio Marcelo de Oliveira Filho <caio.oliveira@intel.com>	2019-11-11 17:17:02 +00:00
Eric Engestrom	86d3a346f1	egl: fix _EGL_NATIVE_PLATFORM fallback When the X11 or Haiku platforms were compiled in, they would bypass the `_EGL_NATIVE_PLATFORM` fallback by always returning themselves instead. Cc: mesa-stable@lists.freedesktop.org Signed-off-by: Eric Engestrom <eric.engestrom@intel.com> Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2019-11-11 17:14:07 +00:00
Ricardo Garcia	20b403aad0	anv: Unify GetDeviceQueue and GetDeviceQueue2 Avoid duplicating some checks and code by making anv_GetDeviceQueue a subcase of anv_GetDeviceQueue2, like radv does. Signed-off-by: Ricardo Garcia <rgarcia@igalia.com> Reviewed-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com> Reviewed-by: Jason Ekstrand <jason@jlekstrand.net>	2019-11-11 16:14:56 +00:00
Alyssa Rosenzweig	5b31182665	panfrost: Select format-specific blending intrinsics If we have an accelerated path for a particular framebuffer format, let's use it to save a bunch of instructions in a blend shader. [Tomeu: Only use the faster intrinsic on >T760] Signed-off-by: Alyssa Rosenzweig <alyssa.rosenzweig@collabora.com> Signed-off-by: Tomeu Vizoso <tomeu.vizoso@collabora.com> Reviewed-by: Tomeu Vizoso <tomeu.vizoso@collabora.com>	2019-11-11 15:23:44 +00:00
Alyssa Rosenzweig	3295edaadf	pan/midgard: Pack load/store masks While most load/store operations on 32-bit/vec4 intriniscally, some are not and have special type-size-dependent semantics for the mask. We need to convert into this native format. Signed-off-by: Alyssa Rosenzweig <alyssa.rosenzweig@collabora.com> Reviewed-by: Tomeu Vizoso <tomeu.vizoso@collabora.com>	2019-11-11 15:23:44 +00:00
Alyssa Rosenzweig	843874c7c3	pan/midgard: Implement nir_intrinsic_load_output_u8_as_fp16_pan We can use the native Midgard ops for this, depending what chip we're on. Signed-off-by: Alyssa Rosenzweig <alyssa.rosenzweig@collabora.com> Reviewed-by: Tomeu Vizoso <tomeu.vizoso@collabora.com>	2019-11-11 15:23:44 +00:00
Alyssa Rosenzweig	5885b64e42	pan/midgard: Identify ld_color_buffer_u8_as_fp16* There are two versions of this opcode, depending what version of the ISA you're using. I'm not sure if there's a semantic difference; I think there might be some slight subtleties but it's too early to know at this stage. Signed-off-by: Alyssa Rosenzweig <alyssa.rosenzweig@collabora.com> Reviewed-by: Tomeu Vizoso <tomeu.vizoso@collabora.com>	2019-11-11 15:23:44 +00:00
Alyssa Rosenzweig	03f73c7fc6	nir: Add load_output_u8_as_fp16_pan intrinsic This is a single opcode, at least on newer Midgard chips. It's easier to have this represented in NIR rather than trying to optimize out the conversions, so let's add the intrinsic. Signed-off-by: Alyssa Rosenzweig <alyssa.rosenzweig@collabora.com> Reviewed-by: Tomeu Vizoso <tomeu.vizoso@collabora.com>	2019-11-11 15:23:44 +00:00
Tomeu Vizoso	ee5321f239	panfrost: Set depth and stencil for SFBD based on the format Signed-off-by: Tomeu Vizoso <tomeu.vizoso@collabora.com> Reviewed-by: Alyssa Rosenzweig <alyssa.rosenzweig@collabora.com>	2019-11-11 15:23:44 +00:00
Erik Faye-Lund	b4d47e21d7	zink: correct depth-stencil format When using packed vulkan-formats on little-endian systems, we need to swap the components for the gallium formats. And since Zink isn't big-endian safe yet, little-endian is the only endianess we care about right now. This fixes a bunch of piglit tests, amongs others: - spec@arb_depth_texture@depth-level-clamp - spec@arb_depth_texture@depthstencil-render-miplevels * d=z24 - spec@arb_depth_texture@fbo-depth-gl_depth_component24-blit - spec@arb_depth_texture@fbo-depth-gl_depth_component24-copypixels - spec@arb_depth_texture@fbo-depth-gl_depth_component24-drawpixels - spec@arb_depth_texture@fbo-depth-gl_depth_component24-readpixels Signed-off-by: Erik Faye-Lund <erik.faye-lund@collabora.com> Fixes: `8d46e35d16` ("zink: introduce opengl over vulkan")	2019-11-11 14:35:53 +00:00
Erik Faye-Lund	d7a6cc8f4a	zink/spirv: add support for nir_op_flrp This fixes the following piglit: spec@ati_fragment_shader@ati_fragment_shader-render-fog Signed-off-by: Erik Faye-Lund <erik.faye-lund@collabora.com>	2019-11-11 14:25:30 +01:00
Chris Wilson	863872e141	egl: Mention if swrast is being forced The system can be disabling HW acceleration unbeknown to the user, leading to a long debug session trying to work out which component is failing. A quick mention that it is the environment override would be very useful. v2: Use more generic "CPU renderer" and so try to avoid jargon. Reviewed-By: Tapani Pälli <tapani.palli@intel.com> Reviewed-by: Emil Velikov <emil.velikov@collabora.com> Reviewed-by: Eric Engestrom <eric.engestrom@intel.com> Acked-by: Martin Peres <martin.peres@linux.intel.com>	2019-11-11 11:52:02 +00:00
Jason Ekstrand	9e440b8d0b	spirv: Sort out the mess that is sampled image This commit makes two major changes. First, we add a second case to OpLoad for sampled images which constructs a vtn_sampled_image and stashes that rather than stashing a pointer to the combined image sampler like we do for bare samplers and images. This should be more in line with how SPIR-V is intended to work and hopefully doesn't cause any weird problems. The second is a rework of vtn_handle_texture to assume that everything has an image but not everything has a sampler. We also add a vtn_fail_if for the case where a texture instructions require a sampler but none is provided. Reviewed-by: Caio Marcelo de Oliveira Filho <caio.oliveira@intel.com>	2019-11-09 15:29:01 +00:00
Jason Ekstrand	9cc4c2c916	spirv: Add a vtn_decorate_pointer helper This helper makes a duplicate copy of the pointer if any new access flags are set at this stage. This way we don't end up propagating access flags further than they actual SPIR-V decorations. In several instances where we create new pointers, we still call the decoration helper directly because no copy is needed. Reviewed-by: Caio Marcelo de Oliveira Filho <caio.oliveira@intel.com>	2019-11-09 15:29:01 +00:00
Jason Ekstrand	4f9688e571	spirv: Remove the type from sampled_image We have types on all vtn_values at this point so there's no reason to carry the redundant type information. Reviewed-by: Caio Marcelo de Oliveira Filho <caio.oliveira@intel.com>	2019-11-09 15:29:01 +00:00
Rob Clark	a3dc975ee7	freedreno/ir3: also track # of nops for shader-db The instruction count is (mostly) a measure of what optimization passes can do, while # of nops is more an indication of how effectively the scheduler is balancing register pressure vs instruction count. So track these independently. (There could be opportunities to rematerialize values to reduce register pressure, swapping some nop's with other alu instructions, so nothing is truely independent.. but it is still useful to break these stats out.) Signed-off-by: Rob Clark <robdclark@chromium.org>	2019-11-09 02:49:15 +00:00
Rob Clark	5f45818673	freedreno/ir3: sync disasm changes from envytools Signed-off-by: Rob Clark <robdclark@chromium.org>	2019-11-09 02:49:15 +00:00
Rob Clark	f3980a8ef7	freedreno/a4xx: fix SP_FS_MRT_REG.HALF_PRECISION Set flag based on actual output reg type. Signed-off-by: Rob Clark <robdclark@chromium.org>	2019-11-09 02:49:15 +00:00
Rob Clark	f0f9ec6882	freedreno/a3xx: fix SP_FS_MRT_REG.HALF_PRECISION We should really be setting this based on the actual output register type. Signed-off-by: Rob Clark <robdclark@chromium.org>	2019-11-09 02:49:15 +00:00
Rob Clark	df229977c3	freedreno/ir3: remove obsolete comment The meta PHI instruction was removed long ago. And fanin/fanout themselves to not contribute actual instructions (at least not by the time you get to sched, they may prevent copy-propagating away a mov) Signed-off-by: Rob Clark <robdclark@chromium.org>	2019-11-09 02:49:15 +00:00
Rob Clark	e804b42fd7	freedreno/ir3/ra: remove ir print after livein/out The IR hasn't changed at this point, so it isn't really adding any value. Signed-off-by: Rob Clark <robdclark@chromium.org>	2019-11-09 02:49:15 +00:00
Rob Clark	8b92052f10	freedreno/ir3/ra: move regs_count==0 check Fold it in to writes_gpr() (since a register that does not reference any registers by definition does not write a register). This lets us avoid having to handle this case in a few other places. Signed-off-by: Rob Clark <robdclark@chromium.org>	2019-11-09 02:49:15 +00:00
Rob Clark	bd21c73d3f	freedreno/ir3: ir3_print tweaks Handle HALF/HIGH flags in all cases, and colorize SSA src notation. Signed-off-by: Rob Clark <robdclark@chromium.org>	2019-11-09 02:49:15 +00:00
Rob Clark	5da10704bb	freedreno/ir3: use SSA flag on dest register too We did this in some places before, but not consistantly. But it will be useful for two-pass RA, to identify which registers have already been assigned. While we are cleaning this up, use __ssa_src() and new __ssa_dst() helper more consistently. (If nothing else, this reduces the # of callers of ir3_reg_create() to audit that we didn't miss something) Signed-off-by: Rob Clark <robdclark@chromium.org>	2019-11-09 02:49:14 +00:00
Rob Clark	8449f6183f	freedreno/ir3: split pre-coloring to it's own function Signed-off-by: Rob Clark <robdclark@chromium.org>	2019-11-09 02:49:14 +00:00
Caio Marcelo de Oliveira Filho	087ecd9ca5	spirv: Don't leak GS initialization to other stages The stage specific fields of shader_info are in an union. We've likely been lucky that this value was either overwritten or ignored by other stages. The recent change in shader_info layout in commit `84a1a2578d` ("compiler: pack shader_info from 160 bytes to 96 bytes") made this issue visible. Fixes: `cf2257069c` ("nir/spirv: Set a default number of invocations for geometry shaders") Reviewed-by: Jason Ekstrand <jason@jlekstrand.net> Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2019-11-08 16:28:21 -08:00
Marek Olšák	84a1a2578d	compiler: pack shader_info from 160 bytes to 96 bytes Reviewed-by: Caio Marcelo de Oliveira Filho <caio.oliveira@intel.com>	2019-11-08 16:54:08 -05:00
Marek Olšák	9950523368	glsl/linker: pass shader_info to analyze_clip_cull_usage directly This will be needed by the next commit. Reviewed-by: Caio Marcelo de Oliveira Filho <caio.oliveira@intel.com>	2019-11-08 16:54:06 -05:00
Marek Olšák	3ef50b023e	radeonsi/nir: fix compute shader crash due to nir_binary == NULL This partially reverts `8b30114dda`. Fixes: `8b30114dda` "radeonsi/nir: call nir_serialize only once per shader"	2019-11-08 16:47:59 -05:00
Marek Olšák	8b30114dda	radeonsi/nir: call nir_serialize only once per shader We were calling it twice. First serialize it, then use it to compute the cache key. Reviewed-by: Timothy Arceri <tarceri@itsqueeze.com>	2019-11-08 15:30:28 -05:00
Marek Olšák	ad56022b0d	util: add blob_finish_get_buffer Reviewed-by: Timothy Arceri <tarceri@itsqueeze.com>	2019-11-08 15:30:28 -05:00
Eric Anholt	b1f38aed84	u_format: Fix swizzle of A1R5G5B5. Found once I started using the generated unpack code from the Mesa side. Fixes: `4bbaac3782` ("gallium: Add some more channel orderings of packed formats.") Reviewed-by: Erik Faye-Lund <erik.faye-lund@collabora.com>	2019-11-08 11:56:02 -08:00
David Stevens	0466239aae	virgl: support emulating planar image sampling Mesa emulates planar format sampling with per-plane samplers. Virgl now supports this by allowing the plane index to be passed when creating a sampler view from a planar image. With this change, mesa now passes that information to virgl. Signed-off-by: David Stevens <stevensd@chromium.org> Reviewed-by: Lepton Wu <lepton@chromium.org>	2019-11-08 17:06:56 +00:00
Krzysztof Raszkowski	084431ce45	gallium/swr: Enable some ARB_gpu_shader5 extensions Enable / add to features.txt: - Enhanced textureGather. - Geometry shader instancing. - Geometry shader multiple streams. Reviewed-by: Jan Zielinski <jan.zielinski@intel.com>	2019-11-08 16:04:47 +00:00
Krzysztof Raszkowski	e5ed9a1b91	gallium/swr: Fix GS invocation issues - Fixed proper setting gl_InvocationID. - Fixed GS vertices output memory overflow. Reviewed-by: Jan Zielinski <jan.zielinski@intel.com>	2019-11-08 14:52:16 +00:00
Timur Kristóf	911a826141	ac: Handle invalid GFX10 format correctly in ac_get_tbuffer_format. It happens that some games try to access a vertex buffer without a valid format. This case was incorrectly handled by ac_get_tbuffer_format which made ACO emit an invalid instruction. Signed-off-by: Timur Kristóf <timur.kristof@gmail.com> Cc: 19.3 <mesa-stable@lists.freedesktop.org> Reviewed-by: Samuel Pitoiset <samuel.pitoiset@gmail.com>	2019-11-08 13:30:30 +01:00
Boris Brezillon	ee82f9f07e	panfrost: Try to evict unused BOs from the cache The panfrost BO cache can only grow since all newly allocated BOs are returned to the cache (unless they've been exported). With the MADVISE ioctl that's not a big issue because the kernel can come and reclaim this memory, but MADVISE will only be available on 5.4 kernels. This means an app can currently allocate a lot memory without ever releasing it, leading to some situations where the OOM-killer kicks in and kills the app (or even worse, kills another process consuming more memory than the GL app) to get some of this memory back. Let's try to limit the amount of BOs we keep in the cache by evicting entries that have not been used for more than one second (if the app stopped allocating BOs of this size, it's likely to not allocate similar BOs in a near future). This solution is based on the VC4/V3D implementation. Signed-off-by: Boris Brezillon <boris.brezillon@collabora.com> Reviewed-by: Alyssa Rosenzweig <alyssa.rosenzweig@collabora.com>	2019-11-08 11:26:47 +01:00
Boris Brezillon	25059cc41f	panfrost: Move BO cache related fields to a sub-struct We will soon introduce an LRU list to evict BOs that have been unused for more than 1 second. Let's first move all BO cache fields to a sub-struct to clarify which fields are used by the BO caching logic. Signed-off-by: Boris Brezillon <boris.brezillon@collabora.com> Reviewed-by: Alyssa Rosenzweig <alyssa.rosenzweig@collabora.com>	2019-11-08 11:26:47 +01:00
Alyssa Rosenzweig	5f768eda43	pan/midgard: Switch base for vertex texturing on T720 There aren't texture pipeline registers anymore; instead, space is shared with work and ldst registers for output and input respectively. We need to shift the base registers to represent this correctly. Signed-off-by: Alyssa Rosenzweig <alyssa.rosenzweig@collabora.com>	2019-11-08 06:45:03 +00:00
Alyssa Rosenzweig	ac14facf7a	pan/midgard: Pass shader stage to disassembler Vertex texturing behaves differently from fragment texturing on some GPUs. Signed-off-by: Alyssa Rosenzweig <alyssa.rosenzweig@collabora.com>	2019-11-08 06:45:03 +00:00
Alyssa Rosenzweig	515941202d	pan/midgard: Disassemble half-steps correctly The meaning of some bits shifts; we need to account for this to print swizzles sanely. Signed-off-by: Alyssa Rosenzweig <alyssa.rosenzweig@collabora.com>	2019-11-08 06:45:03 +00:00
Alyssa Rosenzweig	ec2af6bc97	pan/midgard: Fix printing of half-registers in texture ops We were using old style half-registers; let's update that to be consistent, preparing us for more disassmbler changes in this area. Signed-off-by: Alyssa Rosenzweig <alyssa.rosenzweig@collabora.com>	2019-11-08 06:45:03 +00:00
Kristian H. Kristensen	4a4fad7f40	freedreno/ir3: Use regid() helper when setting up precolor regs Signed-off-by: Kristian H. Kristensen <hoegsberg@google.com> Reviewed-by: Rob Clark <robdclark@gmail.com>	2019-11-07 16:46:21 -08:00
Kristian H. Kristensen	3699a74a43	freedreno/a6xx: Turn on tessellation shaders Wow. Very triangle. So shader. Signed-off-by: Kristian H. Kristensen <hoegsberg@google.com> Acked-by: Eric Anholt <eric@anholt.net> Reviewed-by: Rob Clark <robdclark@gmail.com>	2019-11-07 16:40:27 -08:00

1 2 3 4 5 ...

117485 Commits All Branches Search

117485 Commits

All Branches